ANÁLISIS COMPARATIVO DE LIBRERÍAS PYTHON PARA EXTRACCIÓN DE TEXTO DESDE DOCUMENTOS PDF DIGITALES

Autores/as

DOI:

https://doi.org/10.56519/jw1x1e61

Palabras clave:

extracción de texto, PDF, Python, procesamiento documental, corpus lingüístico, minería de texto

Resumen

La extracción automática de texto desde documentos PDF constituye una tarea fundamental para aplicaciones de minería de texto, recuperación de información, gestión documental y construcción de corpus lingüísticos. El objetivo de esta investigación fue comparar el desempeño de cuatro librerías Python para extracción de texto desde PDF: PyMuPDF, pdfminer.six, pdfplumber y pypdf. El experimento se desarrolló con cuatro documentos PDF digitales de 10, 25, 50 y 100 páginas con estructura de complejidad media: tablas, encabezados repetitivos, gráficos con texto, dos idiomas; estos fueron procesados bajo las mismas condiciones computacionales y mediante una rutina homogénea de extracción, limpieza básica y segmentación por punto. Las métricas evaluadas fueron tiempo de extracción, total de frases, porcentaje de frases válidas, porcentaje de duplicados, escalabilidad temporal y páginas con error. Los resultados mostraron que PyMuPDF alcanzó el mejor rendimiento temporal, con un promedio de 0.24 segundos por documento, mientras que pdfminer.six obtuvo el mayor porcentaje de frases válidas, con 59.18 %. pdfplumber presentó el menor porcentaje promedio de duplicados, con 38.12 %. El análisis estadístico evidenció que las diferencias de tiempo fueron significativas al considerar la transformación logarítmica y el bloqueo por documento, mientras que las diferencias de calidad textual fueron más moderadas. Se concluye que la selección de la librería debe depender del objetivo de uso: velocidad, mayor recuperación textual o menor redundancia.

Descargas

Los datos de descarga aún no están disponibles.

Referencias

Adobe Systems Incorporated. PDF Reference: Adobe Portable Document Format Version 1.7. 6th ed. San Jose: Adobe Systems; 2006.

International Organization for Standardization. ISO 32000-2:2020 Document management - Portable document format - Part 2: PDF 2.0. Geneva: ISO; 2020.

Jurafsky D, Martin JH. Speech and Language Processing. 3rd draft ed. Stanford: Stanford University; 2024.

Manning CD, Raghavan P, Schutze H. Introduction to Information Retrieval. Cambridge: Cambridge University Press; 2008.

PyMuPDF Documentation. PyMuPDF: Python bindings for MuPDF. Available from: https://pymupdf.readthedocs.io/

pdfminer.six Documentation. pdfminer.six: Python PDF parser and analyzer. Available from: https://pdfminersix.readthedocs.io/

Singer JS. pdfplumber Documentation. Available from: https://github.com/jsvine/pdfplumber

pypdf Documentation. pypdf: A pure-python PDF library. Available from: https://pypdf.readthedocs.io/

Adhikari N, Agarwal S. A comparative study of PDF parsing tools across diverse document categories. arXiv. 2024.

Ferguson N, Pennington J, Beghian N, Mohan A, Kiela D, Agrawal S, et al. ExtractBench: A benchmark and evaluation methodology for complex structured extraction. arXiv. 2026.

Lee T, Kim G, Ahn H, Jeong J, Jeong M, Song J. Integrating OCR and LLMs for enhanced document digitization in ERP systems. IEEE; 2024.

Lakatos R, Urban EK, Szabo Z, Pozsga J, Csernai E, Hajdu A. Designing prompts and creating cleaned scientific text for retrieval augmented generation. IEEE; 2024.

Naik YGR, B S, Amith GK. A review on text extraction techniques for degraded historical document images. IEEE; 2024.

Chelliah BJ, Hariharan M, Prakash A, Manoharan B, Senthilselvi A. Harnessing T5 large language model for enhanced PDF text comprehension and Q&A generation. IEEE; 2024.

Yang H, Wei Z, Cheng Z, Zou M, Chen F, Luo W, et al. Research on performance comparison of different Python PDF parsing libraries in analyzing protection setting lists. IEEE; 2025.

Sharmila SP, Tiwari A. PDFInspect: A unified feature extraction framework for malicious document detection. IEEE; 2026.

Böschen I. Evaluation of the extraction of methodological study characteristics with JATSdecoder. Scientific Reports. 2023.

Sasirekha D, Chandra E. Enhanced techniques for PDF image segmentation and text extraction. IEEE; 2012.

Zaryab MA, Ng CR. Optical character recognition for medical records digitization with deep learning. IEEE; 2023.

Bai L, Mulvenna M, Wang Z, Bond R. Clinical entity extraction: comparison between MetaMap, cTAKES, CLAMP and Amazon Comprehend Medical. IEEE; 2021.

Litvak IL, Kostin A, Lashkin F, Maksiyan T, Lagutin S. Comparison of unsupervised metrics for evaluating judicial decision extraction. arXiv. 2025.

Ramadhan G, Mulyana DI, Adrianto S. Optimization of Tesseract OCR for automatic text extraction on Indonesian ID cards. Indonesian Journal of Systems Engineering and Computer Science. 2025.

Ma Z, Huang F, Zhao L, Guo F, Zhai G, Min X. DocIQ: A benchmark dataset and feature fusion network for document image quality assessment. arXiv. 2025.

Wilhelmi L, Bruns C, Schumann M. Enhancing the extraction of GHG emission-reduction targets from sustainability reports using vision language models. MAKE. 2026.

Descargas

Publicado

2026-07-09

Cómo citar

ANÁLISIS COMPARATIVO DE LIBRERÍAS PYTHON PARA EXTRACCIÓN DE TEXTO DESDE DOCUMENTOS PDF DIGITALES. (2026). Revista Científica Multidisciplinaria InvestiGo, 7(20), 299-313. https://doi.org/10.56519/jw1x1e61

Artículos similares

1-10 de 396

También puede Iniciar una búsqueda de similitud avanzada para este artículo.