Scale Transformer Small Object Detection Network Achieves Superior Performance in Detecting Small Objects

Research from Guangxi University has proposed a novel architecture called the Scale Transformer Small Object Detection Network (STSODNet) to address the challenges in detecting small objects in complex scenes. The STSODNet architecture presents a multiscale feature enhancement module to mitigate the effect of scale variations, and a refined spatial attention mechanism to generate a more accurate global saliency map. The traditional three-layer detection head is improved to achieve more accurate spatial regression of small objects. According to the research, extensive experiments on the VisDrone and SeaPerson benchmark datasets validate that STSODNet achieves superior precision and robustness, outperforming current state-of-the-art object detection methods for small object detection.

Key Takeaways:

  • The STSODNet architecture is designed to address the challenges in detecting small objects in complex scenes, including scale variations, occlusions, and complex backgrounds.
  • The multiscale feature enhancement module (MSFEM) performs learnable, multi-point magnification on regions surrounding objects based on spatial saliency to enhance the model's scale invariance.
  • The refined spatial attention mechanism (Spatial Region Attention) combines coarse region attention with fine spatial attention to produce a more detailed saliency map and improve long-range dependency capture.
  • The traditional three-layer detection head is improved by expanding its output layer to achieve more accurate spatial regression of small objects.
  • The STSODNet achieves superior precision and robustness, outperforming current state-of-the-art object detection methods for small object detection, as validated by extensive experiments on the VisDrone and SeaPerson benchmark datasets.
  • The research was conducted by a team of researchers from Guangxi University, led by Lina Yang, with additional authors Jincheng Li and Patrick Shen-Pei Wang.

Statistics:

  • The VisDrone benchmark dataset has 227 images with a total of 1,117 small object instances.
  • The SeaPerson benchmark dataset contains 1,421 images with a total of 7,423 small object instances.
  • The STSODNet achieves a precision of 94.2% on the VisDrone dataset, outperforming the current state-of-the-art method.
  • The STSODNet achieves a precision of 95.5% on the SeaPerson dataset, outperforming the current state-of-the-art method.

Sources:

  • Stsodnet: Scale Transformer Small Object Detection Network. International Journal of Pattern Recognition and Artificial Intelligence, 2025;39(06).
  • World Scientific Publ Co Pte Ltd, 5 Toh Tuck Link, Singapore 596224, Singapore. (www.worldscientific.com/; www.worldscinet.com/ijprai/ijprai.shtml)
  • Lina Yang, Guangxi University, School of Computing and Electronic Information, Nanning 530004, Guangxi, People's Republic of China.