SMART: Semantic Merging Adaptive Regional Transformer for Image Captioning

Jiang, Fengling and Cao, Yujin and Zou, Le and Yang, Erfu and Li, Chenglong and Luo, Chaomin (2026) SMART: Semantic Merging Adaptive Regional Transformer for Image Captioning. IEEE Transactions on Multimedia. ISSN 1520-9210 (https://doi.org/10.1109/TMM.2026.3717485)

[thumbnail of Jiang-etal-2026-SMART-Semantic-Merging-Adaptive-Regional-Transformer-for-Image-Captioning]
Preview
Text. Filename: Jiang-etal-2026-SMART-Semantic-Merging-Adaptive-Regional-Transformer-for-Image-Captioning.pdf
Accepted Author Manuscript
License: Creative Commons Attribution 4.0 logo

Download (14MB)| Preview

Abstract

Image captioning requires integrating accurate object recognition with a deep understanding of semantic relationships to generate natural descriptions. Existing methods primarily rely on grid-based visual features for image understanding, which suffer from semantic fragmentation and fail to form coherent region-level representations. Furthermore, many approaches do not effectively integrate semantic prior information from the textual modality, resulting in coarse and inaccurate descriptions. To address these challenges, we propose the Semantic Merging Adaptive Regional Transformer (SMART) model, which incorporates three novel modules: a Spatial Continuous Region Aggregation Module (SCRAM) to mitigate fragmentation by aggregating grid features into complete object regions; a Semantic-aware Region Enhancement Module (SREM) to align textual semantic priors with these region representations; and an Adaptive Region-to-Grid Feature Fusion Module (ARGFFM) that dynamically selects key regions to enhance feature capacity. Finally, these enhanced features are subsequently leveraged by a transformer-based encoder-decoder architecture to generate semantically robust and accurate captions. Extensive experimental results on the MS COCO dataset establish that the SMART model achieves state of-the-art performance, confirming the efficacy of the SMART.

ORCID iDs

Jiang, Fengling, Cao, Yujin, Zou, Le, Yang, Erfu ORCID logoORCID: https://orcid.org/0000-0003-1813-5950, Li, Chenglong and Luo, Chaomin;