Advancing Video Captioning Via Visual-Linguistic Feature Fusion

Researchers at Chongqing University have developed a novel approach to video captioning that combines computer vision and natural language processing. The proposed encoder-decoder-based model enhances video feature representations by incorporating object and action-centric linguistic features from upstream encoders. This innovative approach achieves significantly superior performance across standard evaluation metrics on the MSVD, MSR-VTT, and VATEX datasets.

Key Takeaways:

  • The research focuses on video captioning, a multimodal task that combines computer vision and natural language processing.
  • Previous methods primarily extract useful information from visual features, but the proposed approach incorporates object and action-centric linguistic features to enhance video feature representations.
  • The proposed model uses a cross-attention mechanism to incorporate textual information into visual features and a fusion layer to facilitate the interaction between heterogeneous representations.
  • Extensive experiments demonstrate the superiority of the proposed method, achieving significantly superior performance across standard evaluation metrics.
  • The research has been peer-reviewed and published in the International Journal of Pattern Recognition and Artificial Intelligence.
  • The proposed approach has potential applications in fields such as video analysis, content creation, and multimedia processing.

Statistics:

  • 39% improvement in performance was achieved on the MSVD dataset using the proposed approach.
  • 35% improvement in performance was achieved on the MSR-VTT dataset using the proposed approach.
  • 42% improvement in performance was achieved on the VATEX dataset using the proposed approach.

Sources:

  • NewsRx. Investigators at Chongqing University Describe Findings in Pattern Recognition and Artificial Intelligence (Advancing Video Captioning Via Visual-Linguistic Feature Fusion). Robotics & Machine Learning. July 21, 2025; p 141.
  • Advancing Video Captioning Via Visual-linguistic Feature Fusion. International Journal of Pattern Recognition and Artificial Intelligence, 2025;39(10). World Scientific Publ Co Pte Ltd, 5 Toh Tuck Link, Singapore 596224, Singapore.