Research on an OCR-Based Identification Model for Low-Quality Cross-Border Logistics Invoices

Authors

  • Junfeng Zhang School of Computer Science and Technology, Zhejiang Sci-Tech University, Hangzhou, China
  • Wei Shen School of Computer Science and Technology, Zhejiang Sci-Tech University, Hangzhou, China
  • Bohao Zheng School of Computer Science and Technology, Zhejiang Sci-Tech University, Hangzhou, China

DOI:

https://doi.org/10.15837/ijccc.2026.5.7060

Keywords:

OCR, DBNet, ACMix, ABINet, InceptionNeXt

Abstract

Due to complex transportation conditions, cross-border logistics invoices often suffer from image blur and irregular layouts, which makes information extraction challenging. To address this issue, this paper proposes an OCR (Optical Character Recognition)-based recognition approach for crossborder logistics invoices. The proposed framework consists of two main stages: text detection and text recognition.

To improve precision and adaptability in both stages, targeted enhancements are introduced. In the text detection stage, the ACMix module is integrated into the backbone of DBNet to strengthen the detection of small and irregular text regions. By combining the local feature modeling capability of convolution with the global context modeling ability of attention, the detection performance on long text and complex backgrounds is effectively improved. In the text recognition stage, InceptionNeXt is adopted to replace ResNet45 in the visual backbone of ABINet, enhancing multiscale feature extraction. By integrating Inception’s multi-scale design with ConvNeXt’s large-kernel optimization strategy, the proposed model achieves more robust feature representation for blurred and multi-scale text in cross-border documents.

A dataset consisting of 700 real-world cross-border logistics document images provided by the Zhejiang Postal Exchange Bureau is constructed, and extensive experiments are conducted on this dataset. Experimental results demonstrate that the proposed method outperforms existing approaches in OCR tasks for cross-border logistics documents, achieving an end-to-end (E2E) precision of 84.72%, a recall of 74.95%, and an Hmean of 79.54%.

References

Wei Shen et al. "End-to-End Information Extraction from Courier Order Images Using a Neural Network Model with Feature Enhancement". In: Applied Sciences 15.2 (2025), p. 698. https://doi.org/10.3390/app15020698

Xinyu Zhou et al. EAST: An Efficient and Accurate Scene Text Detector. 2017. arXiv: 1704. 03155 [cs.CV].

Jonathan Long, Evan Shelhamer, and Trevor Darrell. "Fully Convolutional Networks for Semantic Segmentation". In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 2015, pp. 3431-3440. https://doi.org/10.1109/CVPR.2015.7298965

Yuliang Liu et al. "Single-Shot Arbitrarily-Oriented Scene Text Detection". In: Proceedings of the European Conference on Computer Vision (ECCV). 2019, pp. 455-471.

Minghui Liao, Baoguang Shi, and Xiang Bai. "TextBoxes++: A Single-Shot Oriented Scene Text Detector". In: IEEE Transactions on Image Processing 27.8 (2018), pp. 3676-3690. https://doi.org/10.1109/TIP.2018.2825107

Wei Liu et al. "SSD: Single Shot MultiBox Detector". In: Proceedings of the European Conference on Computer Vision (ECCV). 2016, pp. 21-37. https://doi.org/10.1007/978-3-319-46448-0_2

Youngmin Baek et al. Character Region Awareness for Text Detection. 2019. arXiv: 1904.01941 [cs.CV].

Wenhai Wang et al. Shape Robust Text Detection with Progressive Scale Expansion Network. 2019. arXiv: 1903.12473 [cs.CV].

Wenhai Wang et al. Efficient and Accurate Arbitrary-Shaped Text Detection with Pixel Aggregation Network. 2020. arXiv: 1908.05900 [cs.CV].

Pengyuan Lyu et al. Mask TextSpotter: An End-to-End Trainable Neural Network for Spotting Text with Arbitrary Shapes. 2018. arXiv: 1807.02242 [cs.CV].

Yiqin Zhu et al. Fourier Contour Embedding for Arbitrary-Shaped Text Detection. 2021. arXiv: 2104.10442 [cs.CV]. https://doi.org/10.1109/CVPR46437.2021.00314

Minghui Liao et al. Real-Time Scene Text Detection with Differentiable Binarization. 2019. arXiv: 1911.08947 [cs.CV].

Baoguang Shi, Xiang Bai, and Cong Yao. An End-to-End Trainable Neural Network for Image- Based Sequence Recognition and Its Application to Scene Text Recognition. 2015. arXiv: 1507. 05717 [cs.CV].

Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton. "ImageNet Classification with Deep Convolutional Neural Networks". In: Advances in Neural Information Processing Systems. Vol. 25. 2012, pp. 1097-1105.

Sepp Hochreiter and Jürgen Schmidhuber. "Long Short-Term Memory". In: Neural Computation 9.8 (1997), pp. 1735-1780. https://doi.org/10.1162/neco.1997.9.8.1735

Alex Graves et al. "Connectionist Temporal Classification: Labelling Unsegmented Sequence Data with Recurrent Neural Networks". In: Proceedings of the 23rd International Conference on Machine Learning (ICML). 2006, pp. 369-376. https://doi.org/10.1145/1143844.1143891

Baoguang Shi et al. Robust Scene Text Recognition with Automatic Rectification. 2016. arXiv: 1603.03915 [cs.CV].

Max Jaderberg et al. Spatial Transformer Networks. 2016. arXiv: 1506.02025 [cs.CV].

Baoguang Shi et al. "ASTER: An Attentional Scene Text Recognizer with Flexible Rectification". In: IEEE Transactions on Pattern Analysis and Machine Intelligence 41.9 (2018), pp. 2035-2048. https://doi.org/10.1109/TPAMI.2018.2848939

Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. Neural Machine Translation by Jointly Learning to Align and Translate. 2014. arXiv: 1409.0473 [cs.CL].

Hui Li et al. Show, Attend and Read: A Simple and Strong Baseline for Irregular Text Recognition. 2019. arXiv: 1811.00751 [cs.CV].

Fenfen Sheng, Zhineng Chen, and Bo Xu. "NRTR: A No-Recurrent Sequence-to-Sequence Model for Scene Text Recognition". In: arXiv preprint arXiv:1908.08526 (2019). https://doi.org/10.1109/ICDAR.2019.00130

Zhi Qiao et al. SEED: Semantics Enhanced Encoder-Decoder Framework for Scene Text Recognition. 2020. arXiv: 2005.10977 [cs.CV]. https://doi.org/10.1109/CVPR42600.2020.01354

Shancheng Fang et al. ABINet++: Autonomous, Bidirectional and Iterative Language Modeling for Scene Text Spotting. 2022. arXiv: 2211.10578 [cs.CV].

Xuran Pan et al. On the Integration of Self-Attention and Convolution. 2022. arXiv: 2111.14556 [cs.CV].

Weihao Yu et al. InceptionNeXt: When Inception Meets ConvNeXt. 2024. arXiv: 2303.16900 [cs.CV].

Additional Files

Published

2026-09-01

Most read articles by the same author(s)

Obs.: This plugin requires at least one statistics/report plugin to be enabled. If your statistics plugins provide more than one metric then please also select a main metric on the admin's site settings page and/or on the journal manager's settings pages.