Research on an OCR-Based Identification Model for Low-Quality Cross-Border Logistics Invoices
DOI:
https://doi.org/10.15837/ijccc.2026.5.7060Keywords:
OCR, DBNet, ACMix, ABINet, InceptionNeXtAbstract
Due to complex transportation conditions, cross-border logistics invoices often suffer from image blur and irregular layouts, which makes information extraction challenging. To address this issue, this paper proposes an OCR (Optical Character Recognition)-based recognition approach for crossborder logistics invoices. The proposed framework consists of two main stages: text detection and text recognition.
To improve precision and adaptability in both stages, targeted enhancements are introduced. In the text detection stage, the ACMix module is integrated into the backbone of DBNet to strengthen the detection of small and irregular text regions. By combining the local feature modeling capability of convolution with the global context modeling ability of attention, the detection performance on long text and complex backgrounds is effectively improved. In the text recognition stage, InceptionNeXt is adopted to replace ResNet45 in the visual backbone of ABINet, enhancing multiscale feature extraction. By integrating Inception’s multi-scale design with ConvNeXt’s large-kernel optimization strategy, the proposed model achieves more robust feature representation for blurred and multi-scale text in cross-border documents.
A dataset consisting of 700 real-world cross-border logistics document images provided by the Zhejiang Postal Exchange Bureau is constructed, and extensive experiments are conducted on this dataset. Experimental results demonstrate that the proposed method outperforms existing approaches in OCR tasks for cross-border logistics documents, achieving an end-to-end (E2E) precision of 84.72%, a recall of 74.95%, and an Hmean of 79.54%.
References
Wei Shen et al. "End-to-End Information Extraction from Courier Order Images Using a Neural Network Model with Feature Enhancement". In: Applied Sciences 15.2 (2025), p. 698. https://doi.org/10.3390/app15020698
Xinyu Zhou et al. EAST: An Efficient and Accurate Scene Text Detector. 2017. arXiv: 1704. 03155 [cs.CV].
Jonathan Long, Evan Shelhamer, and Trevor Darrell. "Fully Convolutional Networks for Semantic Segmentation". In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 2015, pp. 3431-3440. https://doi.org/10.1109/CVPR.2015.7298965
Yuliang Liu et al. "Single-Shot Arbitrarily-Oriented Scene Text Detection". In: Proceedings of the European Conference on Computer Vision (ECCV). 2019, pp. 455-471.
Minghui Liao, Baoguang Shi, and Xiang Bai. "TextBoxes++: A Single-Shot Oriented Scene Text Detector". In: IEEE Transactions on Image Processing 27.8 (2018), pp. 3676-3690. https://doi.org/10.1109/TIP.2018.2825107
Wei Liu et al. "SSD: Single Shot MultiBox Detector". In: Proceedings of the European Conference on Computer Vision (ECCV). 2016, pp. 21-37. https://doi.org/10.1007/978-3-319-46448-0_2
Youngmin Baek et al. Character Region Awareness for Text Detection. 2019. arXiv: 1904.01941 [cs.CV].
Wenhai Wang et al. Shape Robust Text Detection with Progressive Scale Expansion Network. 2019. arXiv: 1903.12473 [cs.CV].
Wenhai Wang et al. Efficient and Accurate Arbitrary-Shaped Text Detection with Pixel Aggregation Network. 2020. arXiv: 1908.05900 [cs.CV].
Pengyuan Lyu et al. Mask TextSpotter: An End-to-End Trainable Neural Network for Spotting Text with Arbitrary Shapes. 2018. arXiv: 1807.02242 [cs.CV].
Yiqin Zhu et al. Fourier Contour Embedding for Arbitrary-Shaped Text Detection. 2021. arXiv: 2104.10442 [cs.CV]. https://doi.org/10.1109/CVPR46437.2021.00314
Minghui Liao et al. Real-Time Scene Text Detection with Differentiable Binarization. 2019. arXiv: 1911.08947 [cs.CV].
Baoguang Shi, Xiang Bai, and Cong Yao. An End-to-End Trainable Neural Network for Image- Based Sequence Recognition and Its Application to Scene Text Recognition. 2015. arXiv: 1507. 05717 [cs.CV].
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton. "ImageNet Classification with Deep Convolutional Neural Networks". In: Advances in Neural Information Processing Systems. Vol. 25. 2012, pp. 1097-1105.
Sepp Hochreiter and Jürgen Schmidhuber. "Long Short-Term Memory". In: Neural Computation 9.8 (1997), pp. 1735-1780. https://doi.org/10.1162/neco.1997.9.8.1735
Alex Graves et al. "Connectionist Temporal Classification: Labelling Unsegmented Sequence Data with Recurrent Neural Networks". In: Proceedings of the 23rd International Conference on Machine Learning (ICML). 2006, pp. 369-376. https://doi.org/10.1145/1143844.1143891
Baoguang Shi et al. Robust Scene Text Recognition with Automatic Rectification. 2016. arXiv: 1603.03915 [cs.CV].
Max Jaderberg et al. Spatial Transformer Networks. 2016. arXiv: 1506.02025 [cs.CV].
Baoguang Shi et al. "ASTER: An Attentional Scene Text Recognizer with Flexible Rectification". In: IEEE Transactions on Pattern Analysis and Machine Intelligence 41.9 (2018), pp. 2035-2048. https://doi.org/10.1109/TPAMI.2018.2848939
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. Neural Machine Translation by Jointly Learning to Align and Translate. 2014. arXiv: 1409.0473 [cs.CL].
Hui Li et al. Show, Attend and Read: A Simple and Strong Baseline for Irregular Text Recognition. 2019. arXiv: 1811.00751 [cs.CV].
Fenfen Sheng, Zhineng Chen, and Bo Xu. "NRTR: A No-Recurrent Sequence-to-Sequence Model for Scene Text Recognition". In: arXiv preprint arXiv:1908.08526 (2019). https://doi.org/10.1109/ICDAR.2019.00130
Zhi Qiao et al. SEED: Semantics Enhanced Encoder-Decoder Framework for Scene Text Recognition. 2020. arXiv: 2005.10977 [cs.CV]. https://doi.org/10.1109/CVPR42600.2020.01354
Shancheng Fang et al. ABINet++: Autonomous, Bidirectional and Iterative Language Modeling for Scene Text Spotting. 2022. arXiv: 2211.10578 [cs.CV].
Xuran Pan et al. On the Integration of Self-Attention and Convolution. 2022. arXiv: 2111.14556 [cs.CV].
Weihao Yu et al. InceptionNeXt: When Inception Meets ConvNeXt. 2024. arXiv: 2303.16900 [cs.CV].
Additional Files
Published
Issue
Section
License
Copyright (c) 2026 junfeng zhang, Wei Shen, Bohao Zheng

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.
ONLINE OPEN ACCES: Acces to full text of each article and each issue are allowed for free in respect of Attribution-NonCommercial 4.0 International (CC BY-NC 4.0.
You are free to:
-Share: copy and redistribute the material in any medium or format;
-Adapt: remix, transform, and build upon the material.
The licensor cannot revoke these freedoms as long as you follow the license terms.
DISCLAIMER: The author(s) of each article appearing in International Journal of Computers Communications & Control is/are solely responsible for the content thereof; the publication of an article shall not constitute or be deemed to constitute any representation by the Editors or Agora University Press that the data presented therein are original, correct or sufficient to support the conclusions reached or that the experiment design or methodology is adequate.






