A Statistical Comparison of Multi-Architecture Models for Retention-Period Classification of Chinese Electronic Official Documents

Authors

  • Haoran Zhang School of Artificial Intelligence, Wenzhou Polytechnic University, Wenzhou, Zhejiang 325035, China
  • Lili Shi School of Artificial Intelligence, Wenzhou Polytechnic University, Wenzhou, Zhejiang 325035, China
  • Tuochen Pan Ruian College, Wenzhou Polytechnic University, Wenzhou, Zhejiang 325200, China

DOI:

https://doi.org/10.54097/0bpxr973

Keywords:

Electronic Document Archiving, Text Classification, McNemar Test, Benjamini-Hochberg Correction, Large Language Models, FastText, Statistical Significance

Abstract

Electronic document archiving requires automatic retention-period classification (permanent/30-year/10-year). To evaluate architectures on weakly structured Chinese official texts—texts lacking clear semantic boundaries whose key cues appear as scattered keywords—we benchmark six models across four paradigms: Qwen3 (0.6B/1.7B/4B) with LoRA, BERT-base-Chinese, FastText, and BERT-Embedding-TextCNN (frozen BERT character embeddings + TextCNN). On 2,118 test samples we conduct pairwise McNemar exact binomial tests with Benjamini-Hochberg correction over the 15 comparisons (FDR q=0.05). BERT-Embedding-TextCNN (75.50%) and FastText (75.07%) achieve the highest accuracy, ahead of Qwen3-4B (73.94%), but their Wilson 95% confidence intervals (approximately ±1.9 percentage points) overlap substantially. After correction, only two differences remain significant: BERT-Embedding-TextCNN and FastText each outperform Qwen3-0.6B (adjusted p=0.002 and p=0.006); the other 13 pairwise differences do not reach significance. The accuracies thus form descriptive clusters rather than statistically verified tiers. A leakage-controlled evaluation on the 1,690 test documents that are neither exact nor near-duplicates of training documents preserves this ranking and both significant differences, with accuracies 1.7–2.5 percentage points lower across models. Feature-weight inspection shows that archival keywords (e.g., bidding and state-asset terms, which skew toward permanent retention) are important predictive features for FastText, supporting a keyword-driven explanation of task success. For resource-constrained enterprise deployment, FastText offers near-top accuracy with CPU-only inference at roughly 26,800× the throughput of the 4B LLM under the evaluated hardware configurations.

Downloads

Download data is not yet available.

References

[1] Wang, Q., & Wu, Z. (2020). Integration Framework of Business Systems and Archives Management Systems: Construction and Connotation Analysis. Archives Science Bulletin, (6), 45–53. [In Chinese]

[2] Qu, F. (2024). On Archives Management in Public Institutions. Lantai Inside and Outside, (22), 43–45. [In Chinese]

[3] Han, F. (2024). Research on the Policy Basis and Supply of Data Archiving Based on the 14th Five Year Plan for National Archives Development. Lantai Inside and Outside, (25), 22–24. [In Chinese]

[4] Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017, June 12). Attention Is All You Need. arXiv. https://arxiv.org/abs/1706.03762

[5] General Office of the CPC Central Committee, & General Office of the State Council. (2021, June 9). 14th Five Year Plan for National Archives Development. https://www.saac.gov.cn/daj/toutiao/202106/ecca2de5bce44a0eb55c890762868683.shtml [In Chinese]

[6] Joulin, A., Grave, E., Bojanowski, P., & Mikolov, T. (2017). Bag of Tricks for Efficient Text Classification. In Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics (pp. 427–431).

[7] Devlin, J., Chang, M. W., Lee, K., & Toutanova, K. (2019). BERT: Pre training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics (pp. 4171–4186).

[8] Qwen Team. (2025). Qwen3 Technical Report. Alibaba Group.

[9] McNemar, Q. (1947). Note on the Sampling Error of the Difference Between Correlated Proportions or Percentages. Psychometrika, 12(2), 153–157.

[10] Public Record Office Victoria. (2020). Victorian Government Email Machine Assisted Appraisal: Proof of Concept. https://prov.vic.gov.au/

[11] The National Archives. (2021, November 8). Using AI for Digital Selection in Government. https://www.nationalarchives.gov.uk/

[12] Pistol, R., & Freise, K. (2024). EHRI Innovation Report (D3.4). King’s College London.

[13] European Commission. (2020, June 5). Time Machine Science and Technology Roadmap (D2.2). https://www.timemachine.eu/wp content/uploads/2020/07/D2.2_TM_CSA_Science_and_Technology_P1_Roadmap_v1.5.pdf

[14] Li, X., & Shu, Z. (2021). Research on Intelligent Archiving Method of Electronic Documents Based on Metadata. Archives, (9), 48–52. [In Chinese]

[15] Kang, Y., & Yuan, J. (2023). Application of “Multi Agent” Technology in Electronic Document Archiving Management of “One Network Unified Service” in Government Services. China Archives, (4), 64. [In Chinese]

[16] Chen, X. (2022). Artificial Intelligence and Archives Management: Progress, Vision, and Challenges. China Archives, (11), 30–32. [In Chinese]

[17] Kim, Y. (2014). Convolutional Neural Networks for Sentence Classification. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (pp. 1746–1751).

[18] Hu, J., Shen, L., & Sun, G. (2018). Squeeze and Excitation Networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (pp. 7132–7141).

[19] Yuan, Y., & Zou, Y. (2023). SECNN: A SE Channel Based Convolutional Neural Network for Chinese Text Classification. arXiv. https://arxiv.org/abs/2312.06088

[20] Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert Voss, A., Krueger, T., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., … Amodei, D. (2020). Language Models are Few Shot Learners. In Advances in Neural Information Processing Systems (Vol. 33, pp. 1877–1901).

[21] Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M. A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., Rodriguez, A., Joulin, A., Grave, E., & Lample, G. (2023). LLaMA: Open and Efficient Foundation Language Models. Meta AI.

[22] Zheng, Y. (2024). LLaMA Factory: Unified Efficient Fine Tuning of 100+ Language Models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (pp. 1–9).

[23] Ma, Y. B., Cao, Y. X., Hong, Y. C., Li, Y., & Sun, M. (2023). Large Language Model Is Not a Good Few shot Information Extractor, but a Good Reranker for Hard Samples! In Findings of the Association for Computational Linguistics: EMNLP 2023 (pp. 10572–10601).

[24] Dietterich, T. G. (1998). Approximate Statistical Tests for Comparing Supervised Classification Learning Algorithms. Neural Computation, 10(7), 1895–1923.

[25] Zhang, Y. F., & Ding, M. (2022). Involved Text Recognition Based on Improved Bidirectional Encoder Representation from Transformers. Science Technology and Engineering, 22(29), 12945–12953. [In Chinese]

[26] Wen, W. Z. H. (2022). News Text Classification System Based on Deep Learning [Master’s thesis]. Nanjing University of Posts and Telecommunications. [In Chinese]

[27] Zhang, H., Shan, Y., Peng, J., & Li, C. (2022). A Text Classification Method Based on BERT Att TextCNN Model. In Proceedings of the 2022 IEEE International Conference on Mechanical, Electronic, and Information Technology (IMCEC) (pp. 1731–1735).

[28] National Archives Administration of China. (2021). Regulations on the Administration of Enterprise Archives (Order No. 10). https://www.saac.gov.cn/ [In Chinese]

[29] Benjamini, Y., & Hochberg, Y. (1995). Controlling the False Discovery Rate: A Practical and Powerful Approach to Multiple Testing. Journal of the Royal Statistical Society: Series B (Methodological), 57(1), 289–300.

Downloads

Published

31-08-2026

Issue

Section

Articles