Deep Reinforcement Learning Algorithm Outperforms Traditional Approaches in Multi-Objective Traveling Salesman Problem

Researchers at Shanghai University have made significant advancements in solving the multi-objective traveling salesman problem (MOTSP) by proposing a deep reinforcement learning (DRL) algorithm using a cross fusion attention network (CFAN). According to the study, traditional algorithms often face challenges in efficiently finding satisfactory solutions due to the vast search space and inherent conflicts between objectives. The CFAN architecture, on the other hand, is designed to capture the relationships between problem instances and weight preferences, thereby constructing unified context features. This enables a single trained CFAN model to solve problems with varying weight preferences.

The researchers conducted a comparative analysis with classical evolutionary algorithms and advanced DRL approaches across various MOTSP instances, demonstrating that CFAN consistently outperforms both categories of algorithms, achieving superior solution quality and generalization capability. Specifically, CFAN achieves a 1.43% improvement in the hypervolume (HV) metric over the best-performing DRL algorithm on KroAB instances, a 3.12% improvement on tri-objective problem instances, and a 2.17% improvement on large-scale problem instances.

Key Takeaways:

  • The multi-objective traveling salesman problem (MOTSP) is a classical type of multi-objective combinatorial optimization problem (MOCOP) with numerous real-world applications.
  • Traditional algorithms often face challenges in efficiently finding satisfactory solutions due to the vast search space and inherent conflicts between objectives.
  • The proposed deep reinforcement learning (DRL) algorithm using a cross fusion attention network (CFAN) is designed to capture the relationships between problem instances and weight preferences, enabling a single trained CFAN model to solve problems with varying weight preferences.
  • CFAN consistently outperforms classical evolutionary algorithms and advanced DRL approaches, achieving superior solution quality and generalization capability.
  • CFAN achieves a 1.43% improvement in the hypervolume (HV) metric over the best-performing DRL algorithm on KroAB instances, a 3.12% improvement on tri-objective problem instances, and a 2.17% improvement on large-scale problem instances.
  • The study highlights the effectiveness of CFAN in handling diverse problem instances.

Statistics:

  • 1.43% improvement in the hypervolume (HV) metric over the best-performing DRL algorithm on KroAB instances.
  • 3.12% improvement on tri-objective problem instances.
  • 2.17% improvement on large-scale problem instances.

Sources:

  • Optimizing the multi-objective traveling salesman problem with a deep reinforcement learning algorithm using cross fusion attention networks. Neural Networks, 2025;192:107904.
  • Xiaoyu Fu, School of Mechatronic Engineering and Automation, Shanghai University, 99 Shangda Road, Shanghai, 200444, People's Republic of China.
  • Shenshen Gu and Chee-Meng Chew, authors of the research.
  • Pergamon-elsevier Science Ltd, The Boulevard, Langford Lane, Kidlington, Oxford OX5 1GB, England.