Abstract
OCR-dependent document understanding pipelines remain highly relevant in practical document AI because downstream models explicitly consume OCR text together with bounding box metadata. However, despite extensive studies on textual and image-level attacks, the robustness of bounding box metadata perturbation at the downstream reasoning stage remains insufficiently understood. To overcome this, we introduce LayoutAttack, a systematic benchmark based on Selective Projected Gradient Descent (S-PGD) for perturbing bounding box metadata under architecture-aware layout interfaces. We evaluate five representative OCR-dependent models across four datasets, two task families, two perturbation granularities, and four attack strategies. Under PGD-all at budget 5, LayoutLMv3 shows the largest degradation, reaching 34.52 on DocVQA and 14.15 on SROIE at word level, with averages of 15.68 and 8.54 at word and line level, respectively. In contrast, DocLLM remains comparatively stable (1.57/2.30), while DocLayLLM shows only small average changes (0.51/0.82). These results suggest that vulnerability under bounding box perturbation is associated with effective geometric coupling: models whose predictions remain more tightly tied to OCR coordinates tend to be more attackable. Overall, our findings indicate that robustness depends not simply on whether a model uses layout, but also on how layout information is injected, as different interface designs create systematically different attack surfaces.
Data Availability
No datasets were generated or analysed during the current study.
References
Agrawal, A., Lu, J., Antol, S., Mitchell, M., Zitnick, C. L., Batra, D., & Parikh, D. (2016). VQA: Visual question answering. arxiv:1505.00468.
Bai, S., Chen, K., Liu, X., Wang, J., Ge, W., Song, S., Dang, K., Wang, P., Wang, S., Tang, J., Zhong, H., Zhu, Y., Yang, M., Li, Z., Wan, J., Wang, P., Ding, W., Fu, Z., Xu, Y., Ye, J., Zhang, X., Xie, T., Cheng, Z., Zhang, H., Yang, Z., Xu, H., & Lin, J. (2025). Qwen2.5-VL technical report. arxiv:2502.13923
Bensch, O., Popa, M., & Spille, C. (2021). Key information extraction from documents: Evaluation and generator. arxiv:2106.14624.
Boucher, N., Blessing, J., Shumailov, I., Anderson, R., & Papernot, N. (2023). When vision fails: Text attacks against ViT and OCR.
Chen, Y., Zhang, J., Peng, K., Zheng, J., Liu, R., Torr, P., & Stiefelhagen, R. (2024). Rodla: Benchmarking the robustness of document layout analysis models. CVPR.
Chen, Z., Wu, J., Wang, W., Su, W., Chen, G., Xing, S., & Dai, J. (2024). Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 24185–24198 (2024)
Cloud, G. (2024). Enterprise Document OCR - Google Cloud Document AI. cloud.google.com/document-ai/docs/form-parser (pp. 2025–2030)
Deng, J., Dong, L., Chen, J., Yan, D., Wang, R., Ye, D., Zhao, L., & Tian, J. (2023). Universal defensive underpainting patch: Making your text invisible to optical character recognition. In Proceedings of the 31st ACM international conference on multimedia (pp. 7559–7568)
Ding, Y., Luo, S., Dai, Y., Jiang, Y., Li, Z., Martin, G., & Peng, Y. (2025). A survey on MLLM-based visually rich document understanding: Methods, challenges, and emerging trends. arxiv:2507.09861.
Du, W., Xue, J., Yang, X., Guo, W., Gu, D., & Han, W. (2025). Transfficformer: A novel transformer-based framework to generate evasive malicious traffic. Knowledge-Based Systems, 319, Article 113546. https://doi.org/10.1016/j.knosys.2025.113546
Feng, H., Wei, S., Fei, X., Shi, W., Han, Y., Liao, L., Lu, J., Wu, B., Liu, Q., Lin, C., Tang, J., Liu, H., & Huang, C. (2025) Dolphin: Document image parsing via heterogeneous anchor prompting. arxiv:2505.14059
He, J., Hu, Y., Wang, L., Xu, X., Liu, N., Liu, H., & Shen, H.T. (2023). Do-GOOD: Towards distribution shift evaluation for pre-trained visual document understanding models. arxiv:2306.02623
Huang, Y., Xu, Y., & Cui, L., al. (2022). Layoutlmv3: Pre-training for document ai with unified text and image masking. In Findings of the association for computational linguistics: EMNLP 2022.
Huang, Z., Chen, K., He, J., Bai, X., Karatzas, D., Lu, S., & Jawahar, C. (2019). Icdar 2019 competition on scanned receipt ocr and information extraction. In 2019 International conference on document analysis and recognition (ICDAR) (pp. 1516–1520). IEEE.
Jaume, G., Ekenel, H. K., & Thiran, J.-P. (2019). Funsd: A dataset for form understanding in noisy scanned documents (p. ICDAR)
Kim, G., Hong, T., Yim, M., Nam, J., Park, J., Yim, J., Hwang, W., Yun, S., Han, D., & Park, S. (2022). Ocr-free document understanding transformer. In European conference on computer vision, pp. 498–517. Springer
Li, L., Ma, R., Guo, Q., Xue, X., & Qiu, X. (2020). Bert-attack: Adversarial attack against bert using bert. In Proceedings of the 2020 conference on empirical methods in natural language processing (EMNLP) (pp. 6193–6202)
Li, Z., Liu, Y., Liu, Q., Ma, Z., Zhang, Z., Zhang, S., Yang, B., Guo, Z., Zhang, J., Wang, X., & Bai, X. (2026). MonkeyOCR: Document parsing with a structure-recognition-relation triplet paradigm. arxiv:2506.05218
Liao, W., Wang, J., Li, H., Wang, C., Huang, J., & Jin, L. (2025). Doclayllm: An efficient multi-modal extension of large language models for text-rich document understanding. Proceedings of the IEEE/CVF conference on computer vision and pattern recognition (pp. 4038–4049)
Liu, D., Yang, M., Qu, X., Zhou, P., Cheng, Y., & Hu, W. (2025). A survey of attacks on large vision-language models: Resources, advances, and future trends. IEEE Transactions on Neural Networks and Learning Systems, 36(11), 19525–19545. https://doi.org/10.1109/TNNLS.2025.3592935
Lu, H., Liu, W., Zhang, B., Wang, B., Dong, K., Liu, B., Sun, J., Ren, T., Li, Z., Yang, H., Sun, Y., Deng, C., Xu, H., Xie, Z., & Ruan, C. (2024). DeepSeek-VL: Towards real-world vision-language understanding.
Lu, J., Yu, H., Wang, Y., Ye, Y., Tang, J., Yang, Z., Wu, B., Liu, Q., Feng, H., Wang, H., Liu, H., & Huang, C. (2025). A bounding box is worth one token: Interleaving layout and text in a large language model for document understanding. arxiv:2407.01976
Lu, Z., Xu, N., Tian, H., Wang, L., & Liu, A.-A. (2025). Medical vlp model is vulnerable: Toward multimodal adversarial attack on large medical vision-language models. In IEEE transactions on circuits and systems for video technology. https://doi.org/10.1109/TCSVT.2025.3602970
Luo, C., Cheng, C., Zheng, Q., & Yao, C. (2023). Geolayoutlm: Geometric pre-training for visual information extraction. In IEEE/CVF conference on computer vision and pattern recognition (CVPR).
Luo, C., Shen, Y., Zhu, Z., Zheng, Q., Yu, Z., & Yao, C. (2024). Layoutllm: Layout instruction tuning with large language models for document understanding. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition (pp. 15630–15640)
Mathew, M., Karatzas, D., & Jawahar, C. V. (2021). Docvqa: A dataset for vqa on document images. In Proceedings of the IEEE/CVF winter conference on applications of computer vision (WACV).
Microsoft Azure: Read API — Azure AI Document Intelligence. https://learn.microsoft.com/en-us/azure/ai-services/document-intelligence/prebuilt/layout?view=doc-intel-4.0.0&tabs=rest%2Csample-code Accessed: 2025-05-16 (2024)
Park, S., Shin, S., Lee, B., Lee, J., Surh, J., Seo, M., & Lee, H. (2019). Cord: A consolidated receipt dataset for post-ocr parsing.
Peng, Q., Pan, Y., Wang, W., Luo, B., Zhang, Z., Huang, Z., & Wang, H. (2022) Ernie-layout: Layout knowledge enhanced pre-training for visually-rich document understanding. In Findings of the association for computational linguistics: EMNLP 2022, pp. 3744–3756 (2022)
Piryani, B., Mozafari, J., Abdallah, A., Doucet, A., & Jatowt, A. (2025). Evaluating robustness of LLMs in question answering on multilingual noisy OCR data. arxiv:2502.16781.
Song, C., & Shmatikov, V. (2018). Fooling OCR systems with adversarial text images. arxiv:1802.05385.
Tu, Y., Guo, Y., Chen, H., & Tang, J. (2023). LayoutMask: Enhance text-layout interaction in multi-modal pre-training for document understanding. arxiv:2305.18721.
Wang, D., Raman, N., Sibue, M., Ma, Z., Babkin, P., Kaur, S., Pei, Y., Nourbakhsh, A., & Liu, X. (2024). Docllm: A layout-aware generative language model for multimodal document understanding. In Proceedings of the 62nd annual meeting of the association for computational linguistics (Volume 1: Long Papers), pp. 8529–8548
Wang, H., Zhang, W., Wang, Z., Lu, Z., & Ma, Y. (2026). Vismodal: Visual analytics for evaluating and improving corruption robustness of vision-language models. In IEEE transactions on visualization and computer graphics, pp. 615–625. https://doi.org/10.1109/TVCG.2025.3634257
Wang, Z., Guan, T., Fu, P., Duan, C., Jiang, Q., Guo, Z., Guo, S., Luo, J., Shen, W., & Yang, X. (2025). Marten: Visual question answering with mask generation for multi-modal document understanding. arxiv:2503.14140
Xu, Y., Li, M., Cui, L., Huang, S., Wei, F., & Zhou, M. (2020). Layoutlm: Pre-training of text and layout for document image understanding. In Proceedings of the 26th ACM SIGKDD international conference on knowledge discovery & data mining.
Xu, Y., Xu, Y., Lv, T., Cui, L., Wei, F., Wang, G., Lu, Y., Florencio, D., Zhang, C., Che, W., Zhang, M., & Zhou, L. (2022). LayoutLMv2: Multi-modal pre-training for visually-rich document. arxiv:2012.14740
Ye, J., Hu, A., Xu, H., Ye, Q., Yan, M., Dan, Y., Zhao, C., Xu, G., Li, C., Tian, J., Qi, Q., Zhang, J., & Huang, F. (2023). mPLUG-DocOwl: modularized multimodal large language model for document. arxiv:2307.02499
Ye, J., Hu, A., Xu, H., Ye, Q., Yan, M., Xu, G., Li, C., Tian, J., Qian, Q., Zhang, J., Jin, Q., He, L., Lin, X.A., & Huang, F. (2023). UReader: Universal OCR-free visually-situated language understanding with multimodal large language. arxiv:2310.05126
Yuke, Z., Yue, Z., Dongdong, L., Chi, X., Zihu, X., Bo, Z., & Sheng, G. (2025). Enhancing document understanding with group position embedding: A novel approach to incorporate layout information.
Zhang, C., Zhou, L., Xu, X., Wu, J., & Liu, Z. (2025). Adversarial attacks of vision tasks in the past 10 years: A survey. ACM Computing Surveys, 58(2), 1–42. https://doi.org/10.1145/3743126
Zhu, Z., Luo, C., Shao, Z., Gao, F., Xing, H., Zheng, Q., & Zhang, J. (2025). A simple yet effective layout token in large language models for document understanding. arxiv:2503.18434.
Author information
Authors and Affiliations
Contributions
D.N.T and L.D.D conceptualize the research ideas and implement the empirical evaluation. V.H.V, N.V.T, and M.N.A contribute towards the evaluation of the framework. All authors contributed and reviewed the manuscript.
Corresponding author
Ethics declarations
Conflict of interest
The authors declare no conflict of interest.
Additional information
Editors: Lan Du, Benjamin C. M. Fung, Hady W. Lauw, Longbing Cao.
Publisher's Note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Appendices
Appendix A: Detailed Experimental Setup
This appendix provides implementation and evaluation details for the experiments reported in the main paper, including the hardware environment, benchmark datasets, model configurations, attack procedure, and evaluation metrics.
1.1 Hardware and Software Environment
All experiments are conducted on a single Linux server node with the following specifications:
-
GPUs: 8\(\times\) NVIDIA L40S, each with 48 GB GDDR6 VRAM, connected via PCIe Gen4; NVIDIA driver version 560.35.03 and CUDA toolkit version 12.6.
-
CPUs: 2\(\times\) AMD EPYC 9554 64-Core Processor (128 physical cores, 256 hardware threads in total).
-
System memory: 503 GiB DDR5 RAM.
-
Operating system: Ubuntu-based Linux (kernel 6.2.0-1015-nvidia).
The principal software dependencies are listed in Table 6.
DocLayLLM is loaded in half precision (float16) with automatic device placement (device_map="auto") to fit within the available GPU memory budget. LayoutLMv3 fine-tuning and its attack computations are performed in full precision (float32). Unless otherwise stated, all training runs use seed 42 and all attack runs use seed 2025.
1.2 Benchmark Datasets
We evaluate on four widely used document-understanding benchmarks spanning two task families: key information extraction (KIE) and document visual question answering (DocVQA). Table 7 summarizes the dataset statistics.
KIE datasets. FUNSD consists of noisy scanned forms with four entity types (header, question, answer, and other). CORD contains Indonesian receipts with 30 fine-grained field types. SROIE provides scanned receipts with four key fields (company, date, address, and total).
DocVQA. DocVQA pairs document images with natural-language questions and one or more reference answers. OCR annotations are taken from the Azure Read API outputs distributed with the dataset. We evaluate on the full validation split.
OCR granularity. Each dataset is processed at two OCR granularities:
-
Word-level: each OCR-recognized word is assigned its own bounding box \([x_1, y_1, x_2, y_2]\).
-
Line-level: all words on the same OCR line share a single enclosing bounding box.
Both granularities are evaluated independently.
1.3 Target Models and Evaluation Protocol
Pretrained checkpoints. We use publicly available pretrained checkpoints from Hugging Face whenever applicable: LayoutLMv3 (microsoft/layoutlmv3-base), ERNIE-Layout (PaddlePaddle/ernie-layoutx-base-uncased), LayTextLLM (LayTextLLM/LayTextLLM-Zero) and DocLLM (JinghuiLuAstronaut/DocLLM_baichuan2_7b).
For reproducibility, the corresponding model pages are: microsoft/layoutlmv3-base, https://www.PaddlePaddle/ernie-layoutx-base-uncased LayTextLLM/LayTextLLM-Zero and DocLLM/JinghuiLuAstronaut/DocLLM_baichuan2_7b.
We report results for five document-understanding models: LayoutLMv3, ERNIE-Layout, LayTextLLM, DocLLM, and DocLayLLM. These models cover both encoder-style discriminative architectures and decoder-style generative architectures.
Discriminative encoders. For LayoutLMv3 and ERNIE-Layout, we fine-tune on the clean training split and then attack the frozen checkpoint at inference time. LayoutLMv3 is used for both KIE and DocVQA, whereas ERNIE-Layout is evaluated on KIE benchmarks only.
Generative LLM-based models. For LayTextLLM, DocLLM, and DocLayLLM, we use the released checkpoints in zero-shot inference mode and inject perturbations directly into the OCR-derived layout interface. These models generate textual outputs autoregressively.
Primary implementation focus. The most detailed attack implementation in this appendix is given for LayoutLMv3 and DocLayLLM, which serve as the representative encoder-based and decoder-based models in the main paper. The same attack principles are then applied consistently across the other evaluated models wherever supported by their layout input interface.
1.4 Fine-Tuning Protocols
1.4.1 LayoutLMv3 for KIE
For each of the three KIE datasets, we independently fine-tune the pre-trained layoutlmv3-base checkpoint with a token classification head. Fine-tuning is performed separately for word-level and line-level inputs.
The complete training configuration is listed in Table 8. We use a linear learning-rate schedule with a short warmup phase. Early stopping monitors validation entity-level F1 and halts training if no improvement is observed for 15 consecutive evaluations. The best checkpoint is retained for attack-time evaluation.
1.4.2 LayoutLMv3 for DocVQA
We fine-tune LayoutLMv3 with an extractive question-answering head on the DocVQA training split. The input is formatted as a question-context pair: the question is tokenized first, followed by the OCR tokens and their bounding boxes. The model predicts start and end positions within the context span.
Due to the larger training set, we use multi-GPU distributed training with PyTorch torchrun across two GPUs. Early stopping monitors validation ANLS and halts training after 20 evaluations without improvement. The full configuration is given in Table 9.
1.4.3 ERNIE-Layout for KIE
For the KIE benchmarks, we fine-tune the pre-trained ernie-layoutx-base-uncased checkpoint separately on FUNSD, CORD, and SROIE using a token classification head with BIO tagging. Fine-tuning is performed independently for word-level and line-level inputs.
To keep the KIE evaluation protocol uniform for ERNIE-Layout, we use the same training recipe across all three datasets rather than mixing a released setup for one dataset with newly trained models for the others. The complete training configuration is listed in Table 10. We evaluate and save checkpoints periodically during training, retain only the best checkpoint according to validation F1, and use that checkpoint for attack-time evaluation.
For DocVQA, we use the publicly released ERNIE-Layout checkpoint associated with Chinese DocVQA rather than introducing an additional task-specific training recipe for this benchmark. In our pipeline, this Chinese DocVQA setup is supported only at the line level; therefore, no word-level DocVQA result is reported for ERNIE-Layout.
1.5 Attack Methodology
Our adversarial attack operates exclusively on the layout channel: it perturbs the bounding-box coordinates \([x_1, y_1, x_2, y_2]\) associated with OCR units, while keeping the textual content and document image pixels unchanged. All coordinates are normalized to [0, 1000].
Coordinate constraints. Every perturbation is projected back into a feasible set that preserves valid box geometry. Coordinates are clamped to [0, 1000], the ordering constraints \(x_1 < x_2\) and \(y_1 < y_2\) are enforced, each coordinate shift is bounded by a maximum perturbation budget \(\epsilon _{\max }\), and the perturbed box area must remain above a minimum ratio of a valid side length \(r_{\min }\) relative to the original box.
1.5.1 Box-Selection Strategies
To study the role of box importance, we evaluate four attack strategies:
Selective vulnerable PGD. We estimate a vulnerability score for each unique bounding box using a leave-one-out ablation procedure. Boxes are ranked by the resulting performance drop, and PGD is applied only to the top-k selected boxes.
Random shift. We use the same selected boxes as in Vulnerable, but replace PGD with random geometric translation. Each selected box is shifted in one of four cardinal directions by a random offset \(\delta \sim \mathcal {U}(0,\epsilon _{\max })\), subject to an IoU acceptance threshold \(\theta _{\textrm{IoU}}\).
Random selective PGD. We select a fixed fraction of boxes uniformly at random and apply PGD to that subset. This isolates the contribution of vulnerability ranking from the optimization procedure itself.
PGD-all. PGD is applied to all bounding boxes simultaneously, without any selection mask.
1.5.2 PGD Optimization Loop
Given a selected set of boxes, PGD proceeds as follows:
-
1.
Initialization. The adversarial boxes are initialized from the clean coordinates.
-
2.
Gradient computation. The model is run with the relaxed layout interface, and the task loss is back-propagated to the box coordinates.
-
3.
Update. Coordinates are updated along the sign of the gradient with step size \(\alpha\), restricted by the selection mask.
-
4.
Projection. After each update, coordinates are projected back into the feasible set described above, while enforcing the minimum ratio of a valid side length compared to original box \(r_{\min }\).
-
5.
Rounding. For models that consume integer coordinates, the continuous values are periodically snapped to the nearest integer.
-
6.
Early stopping. For LayoutLMv3, the best adversarial configuration is tracked on the original discrete model and optimization stops early if the objective does not improve.
-
7.
Restoration. The original discrete embedding tables are restored and the final adversarial coordinates are evaluated on the unmodified model.
1.6 Attack Hyperparameters
Table 11 summarizes the main attack hyperparameters for the representative model families used in our implementation: LayoutLMv3, DocLayLLM and LayTextLLM. These settings differ because the generative models require longer optimization with smaller per-step updates.
Appendix B: Sensitivity Analysis of the Selective Ratio \(\rho\)
Table 12 reports representative selective-ratio sensitivity results used to support the discussion in Sect. 6.2. We evaluate \(\rho \in \{0.25, 0.50, 0.75\}\) and compare them against Random shift, Random selective PGD, and PGD-all on representative OCR-dependent architectures, including LayoutLMv3, LayTextLLM, and DocLayLLM, on FUNSD and CORD under both word-level and line-level perturbations.
Overall, increasing \(\rho\) generally increases attack strength, but the main robustness ordering remains stable. LayoutLMv3 is consistently more sensitive to bbox perturbation, LayTextLLM exhibits moderate sensitivity with stronger degradation on CORD, and DocLayLLM remains comparatively stable. These results support the use of \(\rho =0.5\) as the default setting in the main benchmark while showing that the core selective-attack conclusions are not artifacts of a single ratio choice.
Rights and permissions
Springer Nature or its licensor (e.g. a society or other partner) holds exclusive rights to this article under a publishing agreement with the author(s) or other rightsholder(s); author self-archiving of the accepted manuscript version of this article is solely governed by the terms of such publishing agreement and applicable law.
About this article
Cite this article
Tien, D.N., Le, D.D., Hoang, D.N. et al. Probing Spatial Robustness in OCR-Dependent Document Understanding: Adversarial Attacks on Bounding Box Metadata Across Layout-Aware Architectures. Mach Learn 115, 203 (2026). https://doi.org/10.1007/s10994-026-07087-x
Received:
Revised:
Accepted:
Published:
Version of record:
DOI: https://doi.org/10.1007/s10994-026-07087-x
Facts Only
* LayoutAttack is a benchmark based on Selective Projected Gradient Descent (S-PGD) for perturbing bounding box metadata under architecture-aware layout interfaces.
* Five document-understanding models were evaluated: LayoutLMv3, ERNIE-Layout, LayTextLLM, DocLLM, and DocLayLLM.
* Evaluation covered four datasets and two task families: key information extraction (KIE) and document visual question answering (DocVQA).
* OCR granularity was tested at word-level and line-level.
* LayoutLMv3 showed the largest degradation under PGD-all attack, reaching 34.52 on DocVQA word level and 14.15 on SROIE at word level.
* DocLLM remained comparatively stable (1.57/2.30).
* The attack operates by perturbing bounding box coordinates while keeping text and image pixels unchanged.
* Attack strategies included Selective vulnerable PGD, Random shift, Random selective PGD, and PGD-all.
* Results suggest vulnerability is associated with effective geometric coupling of model predictions to OCR coordinates.
