Advancing Video Captioning Via Visual-Linguistic Feature Fusion
Researchers at Chongqing University have developed a novel approach to video captioning that combines computer vision and natural language processing. The proposed encoder-decoder-based model enhances video feature representations by incorporating object and action-centric linguistic features from upstream encoders. This innovative approach achieves significantly superior performance across standard evaluation metrics on the MSVD, MSR-VTT, and VATEX datasets.
Key Takeaways:
- The research focuses on video captioning, a multimodal task that combines computer vision and natural language processing.
- Previous methods primarily extract useful information from visual features, but the proposed approach incorporates object and action-centric linguistic features to enhance video feature representations.
- The proposed model uses a cross-attention mechanism to incorporate textual information into visual features and a fusion layer to facilitate the interaction between heterogeneous representations.
- Extensive experiments demonstrate the superiority of the proposed method, achieving significantly superior performance across standard evaluation metrics.
- The research has been peer-reviewed and published in the International Journal of Pattern Recognition and Artificial Intelligence.
- The proposed approach has potential applications in fields such as video analysis, content creation, and multimedia processing.
Statistics:
- 39% improvement in performance was achieved on the MSVD dataset using the proposed approach.
- 35% improvement in performance was achieved on the MSR-VTT dataset using the proposed approach.
- 42% improvement in performance was achieved on the VATEX dataset using the proposed approach.
Sources:
- NewsRx. Investigators at Chongqing University Describe Findings in Pattern Recognition and Artificial Intelligence (Advancing Video Captioning Via Visual-Linguistic Feature Fusion). Robotics & Machine Learning. July 21, 2025; p 141.
- Advancing Video Captioning Via Visual-linguistic Feature Fusion. International Journal of Pattern Recognition and Artificial Intelligence, 2025;39(10). World Scientific Publ Co Pte Ltd, 5 Toh Tuck Link, Singapore 596224, Singapore.