主办:陕西省汽车工程学会
ISSN 1671-7988  CN 61-1394/TH
创刊:1976年

Automobile Applied Technology ›› 2026, Vol. 51 ›› Issue (17): 34-39.DOI: 10.16638/j.cnki.1671-7988.2026.017.005

• Intelligent Connected Vehicle • Previous Articles    

With the advancement and widespread adoption of assisted driving and autonomous driving technologies, traffic sign recognition has become a pivotal component of intelligent transportation systems. To achieve precise identification of traffic sign images by vehicles, this paper proposes a two-stage recognition method based on YOLOv8n and Vision Transformer (ViT). During the object detection phase, the YOLOv8n model combined with an anchor-free bounding box mechanism enables rapid localization of target regions. In the recognition phase, the ViT network is employed to extract global self-attention features from candidate regions, enhancing discrimination accuracy for similar signs. Additionally, image enhancement techniques such as MixUp, CutMix, and Albumentations are applied during training to improve the model's robustness under real-world driving conditions including occlusion and lighting variations. Finally, the effectiveness and robustness of the proposed method are validated on the German traffic sign detection benchmark (GTSDB) and the German traffic sign recognition benchmark (GTSRB).

YANG Jing, LUO Yuan   

  1. Shanxi Vocational University of Engineering Science and Technology
  • Published:2026-09-07
  • Contact: YANG Jing

基于 YOLOv8n 和 Vision Transformer 的 交通标志识别研究

杨晶,罗元   

  1. 山西工程科技职业大学
  • 通讯作者: 杨晶
  • 作者简介:杨晶(1996-),女,硕士,助教,研究方向为人工智能技术在交通领域的应用
  • 基金资助:
    山西工程科技职业大学 2024 年度校科研基金计划项目(KJ202428)

Abstract: With the advancement and widespread adoption of assisted driving and autonomous driving technologies, traffic sign recognition has become a pivotal component of intelligent transportation systems. To achieve precise identification of traffic sign images by vehicles, this paper proposes a two-stage recognition method based on YOLOv8n and Vision Transformer (ViT). During the object detection phase, the YOLOv8n model combined with an anchor-free bounding box mechanism enables rapid localization of target regions. In the recognition phase, the ViT network is employed to extract global self-attention features from candidate regions, enhancing discrimination accuracy for similar signs. Additionally, image enhancement techniques such as MixUp, CutMix, and Albumentations are applied during training to improve the model's robustness under real-world driving conditions including occlusion and lighting variations. Finally, the effectiveness and robustness of the proposed method are validated on the German traffic sign detection benchmark (GTSDB) and the German traffic sign recognition benchmark (GTSRB).

Key words: traffic sign recognition; YOLOv8; intelligent transportation; data augmentation

摘要: 随着辅助驾驶技术及自动驾驶技术的发展与普及,交通标志识别技术已成为智能交通 系统的关键技术之一。为实现车辆对交通标志图像的精准识别,文章提出了一种基于 YOLOv8n 和 Vision Transformer(ViT)的两阶段交通标志识别方法。在目标检测阶段,采用 YOLOv8n 模型结合无锚框机制实现交通标志区域的快速定位;在交通标志识别阶段,引入 ViT 网络对候选区域进行全局自注意力特征提取,提升对相似标志的精细区分能力;同时, 在训练过程中使用 MixUp、CutMix 及 Albumentations 技术进行图像增强,增强模型在遮挡、 光照变化等真实驾驶环境下的鲁棒性。最后,在德国交通标志检测数据集(GTSDB)与德国 交通标志识别数据集(GTSRB)上验证了所提方法的有效性和鲁棒性。

关键词: 交通标志识别;YOLOv8;智能交通;数据增强