[1]PENG Yang,WANG Dejun,MENG Bo,et al.Historical document layout analysis[J].CAAI Transactions on Intelligent Systems,2026,21(3):727-738.[doi:10.11992/tis.202501011]
Copy

Historical document layout analysis

References:
[1] 王军, 杨海峥, 刘石, 等. 系列笔谈之一: 智能时代古典文献学的机遇与挑战[J]. 数字人文, 2022(2): 108-132 WANG Jun, YANG Haizheng, LIU Shi, et al. The first of a series of pen talks: opportunities and challenges of classical philology in the intelligent age[J]. Digital humanities, 2022(2): 108-132
[2] 刘成林, 金连文, 白翔, 等. 文档智能分析与识别前沿: 回顾与展望[J]. 中国图象图形学报, 2023, 28(8): 2223-2252 LIU Chenglin, JIN Lianwen, BAI Xiang, et al. Frontiers of intelligent document analysis and recognition: review and prospects[J]. Journal of image and graphics, 2023, 28(8): 2223-2252
[3] TANG Y Y, LEE S W, SUEN C Y. Automatic document processing: a survey[J]. Pattern recognition, 1996, 29(12): 1931-1952
[4] GAO Liangcai, YI Xiaohan, JIANG Zhuoren, et al. ICDAR2017 competition on page object detection[C]//2017 14th IAPR International Conference on Document Analysis and Recognition. Kyoto: IEEE, 2017: 1417-1422.
[5] 王尚荣. 文档布局分析的多模态学习方法研究与实现[D]. 北京: 北京邮电大学, 2023. WANG Shangrong. Research and implementation of multimodal learning method for document layout analysis[D]. Beijing: Beijing University of Posts and Telecommunications, 2023.
[6] SASSIOUI A, BENOUINI R, EL OUARGUI Y, et al. Visually-rich document understanding: concepts, taxonomy and challenges[C]//2023 10th International Conference on Wireless Networks and Mobile Communications. Istanbul: IEEE, 2023: 1-7.
[7] AFZAL M Z, K?LSCH A, AHMED S, et al. Cutting the error by half: investigation of very deep CNN and advanced training strategies for document image classification[C]//2017 14th IAPR International Conference on Document Analysis and Recognition. Kyoto: IEEE, 2017: 883-888.
[8] DAVIS B, MORSE B, PRICE B, et al. End-to-end document recognition and understanding with dessurt[M]//Computer Vision–ECCV 2022 Workshops. Cham: Springer Nature Switzerland, 2023: 280-296.
[9] BAKKALI S, MING Zuheng, COUSTATY M, et al. VLCDoC: vision-language contrastive pre-training model for cross-modal document classification[J]. Pattern recognition, 2023, 139: 109419
[10] HONG T, KIM D, JI M, et al. BROS: a pre-trained language model focusing on text and layout for better key information extraction from documents[EB/OL]. (2021-09-10)[2024-01-01]. https://arxiv.org/abs/2108.04539.
[11] HUANG Yupan, LYUengchao, CUI Lei, et al. LayoutLMv3: pre-training for document AI with unified text and image masking[C]//Proceedings of the 30th ACM International Conference on Multimedia. Lisboa: ACM, 2022: 4083-4091.
[12] LI Xin, ZHENG Yan, HU Yiqing, et al. Relational representation learning in visually-rich documents[C]//Proceedings of the 30th ACM International Conference on Multimedia. Lisboa: ACM, 2022: 4614-4624.
[13] ZHANG Chong, TU Yi ZHAO Yixi, et al. Modeling layout reading order as ordering relations for visually-rich document Uunderstanding[EB/OL]. (2024-09-29)[2025-01-01]. https://arxiv.org/abs/2409.19672.
[14] AIELLO M, SMEULDERS A M W. Bidimensional relations for reading order detection[M]//EPRINTS-BOOK-TITLE. University of Groningen, Johann Bernoulli Institute for Mathematics and Computer Science, 2003.
[15] FERILLI S, GRIECO D, REDAVID D, et al. Abstract argumentation for reading order detection[C]//Proceedings of the 2014 ACM Symposium on Document Engineering. Fort Collins: ACM, 2014: 45-48.
[16] LI Liangcheng, GAO Feiyu, BU Jiajun, et al. An end-to-end OCR text re-organization sequence learning for rich-text detail image comprehension[M]//Computer Vision – ECCV 2020. Cham: Springer International Publishing, 2020: 85-100.
[17] QUIR?S L, VIDAL E. Reading order detection on handwritten documents[J]. Neural computing and applications, 2022, 34(12): 9593-9611
[18] 马伟洪. 面向古籍文档分析的文字检测识别与阅读顺序理解[D]. 广州: 华南理工大学, 2022. MA Weihong. Character detection, recognition and reading order comprehension for historical document analysis [D]. Guangzhou: South China University of Technology, 2022.
[19] ZHANG Chong, GUO Ya, TU Yi, et al. Reading order matters: Information extraction from visually-rich documents by token path prediction[EB/OL]. (2023-10-17)[2025-01-01]. https://arxiv.org/abs/2310.11016.
[20] QIAO Liang, LI Can, CHENG Zhanzhan, et al. Reading order detection in visually-rich documents with multi-modal layout-aware relation prediction[J]. Pattern recognition, 2024, 150: 110314
[21] 温绍杰, 吴瑞刚, 冯超文, 等. 基于Transformer的多模态级联文档布局分析网络[J]. 浙江大学学报(工学版), 2024, 58(2): 317-324, 369 WEN Shaojie, WU Ruigang, FENG Chaowen, et al. Multimodal cascaded document layout analysis network based on Transformer[J]. Journal of Zhejiang University (engineering science), 2024, 58(2): 317-324, 369
[22] WANG Jiawei, HU Kai, ZHONG Zhuoyao, et al. Detect-order-construct: a tree construction based approach for hierarchical document structure analysis[J]. Pattern recognit, 2017, 156: 110836
[23] REIMERS N, GUREVYCH I. SENTENCe-BERT: sentence embeddings using siamese BERT-Networks[EB/OL]. (2019-08-27)[2025-01-01]. https://arxiv.org/abs/1908.10084.
[24] XU Canhui, LI Yuteng, SHI Cao, et al. HiM: hierarchical multimodal network for document layout analysis[J]. Applied intelligence, 2023, 53(20): 24314-24326
[25] WANG Dongsheng, MA Zhiqiang, NOURBAKHSH A, et al. DocGraphLM: documental graph language model for information extraction[C]//Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval. Taipei: ACM, 2023: 1944-1948.
[26] 徐捷, 邵玉斌, 杜庆治, 等. 结合混合特征提取与深度学习的长文本语义相似度计算[J]. 计算机工程与科学, 2024, 46(8): 1513-1520 XU Jie, SHAO Yubin, DU Qingzhi, et al. Long text semantic similarity calculation combining hybrid feature extraction and deep learning[J]. Computer engineering & science, 2024, 46(8): 1513-1520
[27] PAPINENI K. BLEU: a method for automatic evaluation of MT[R]. New York: IBM, 2001.
[28] HE Kaiming, GKIOXARI G, DOLL?R P, et al. Mask R-CNN[C]//2017 IEEE International Conference on Computer Vision. Venice: IEEE, 2017: 2980-2988.
[29] HUANG Yupan, LYU ngchao, CUI Lei, et al. LayoutLMv3: pre-training for document AI with unified text and image masking[C]//Proceedings of the 30th ACM International Conference on Multimedia. Lisboa: ACM, 2022: 4083-4091.
[30] WANG Z, XU Y, CUI L, et al. Layoutreader: Pre-training of text and layout for reading order detection[EB/OL]. (2021-08-26)[2025-01-01]. https://arxiv.org/abs/2108.11591.
Similar References:

Memo

-

Last Update: 1900-01-01

Copyright © CAAI Transactions on Intelligent Systems