使用PDFMiner转换PDF时部分文档无单词间距问题求助
Hey there! I’ve dealt with this exact frustrating issue before—those run-together words happen because some PDFs (like the academic paper you linked) don’t use actual space characters to separate words. Instead, they rely on adjusting character positions to create visual spacing, which PDFMiner’s default settings miss. Here’s how to fix it:
Why This Happens
Older or layout-focused PDFs (common in academic publishing) often skip literal spaces and use horizontal gaps between characters to mimic word separation. PDFMiner’s default LAParams (Layout Analysis Parameters) are tuned for standard PDFs, so it doesn’t recognize these gaps as spaces.
Modified Conversion Code
We’ll tweak the LAParams settings to tell PDFMiner to detect those visual gaps as word separators. Here’s the updated function, plus your existing download code:
from pdfminer.pdfinterp import PDFResourceManager, PDFPageInterpreter from pdfminer.converter import TextConverter from pdfminer.layout import LAParams from pdfminer.pdfpage import PDFPage from io import StringIO import requests def convert_pdf_to_txt(path): rsrcmgr = PDFResourceManager() retstr = StringIO() codec = 'utf-8' # Key fix: Adjust layout analysis parameters to detect visual word gaps laparams = LAParams( char_margin=1.0, # Lower threshold: gaps larger than this = new word word_margin=0.5, # Ensure small gaps between words aren't merged detect_vertical=True, # Handle any vertical text edge cases line_overlap=0.5 ) device = TextConverter(rsrcmgr, retstr, codec=codec, laparams=laparams) fp = open(path, 'rb') interpreter = PDFPageInterpreter(rsrcmgr, device) password = "" maxpages = 0 caching = True pagenos = set() for page in PDFPage.get_pages(fp, pagenos, maxpages=maxpages, password=password, caching=caching, check_extractable=True): interpreter.process_page(page) fp.close() device.close() text = retstr.getvalue() retstr.close() # Optional: Clean up messy newlines and extra spaces text = '\n'.join([line.strip() for line in text.split('\n') if line.strip()]) return text # Your working PDF download code url = 'http://www.ece.rochester.edu/~gsharma/papers/LocalImageRegisterEI2005.pdf' file_name = './LocalImageRegisterEI2005.pdf' response = requests.get(url) with open(file_name, 'wb') as f: f.write(response.content) # Test the conversion converted_text = convert_pdf_to_txt(file_name) print(converted_text[:500]) # Print first 500 chars to verify
What Each Parameter Does
- char_margin: Default is 2.0. By lowering it to 1.0, we tell PDFMiner that any gap between characters larger than 1.0 units should be treated as a word separator (inserting a space). This is the main fix for your issue.
- word_margin: Adjusts how close two words can be before being merged into one. 0.5 is a safe middle ground for most academic PDFs.
- detect_vertical: Ensures we don’t miss text laid out vertically, which can cause weird merging in some documents.
- The optional cleanup step removes redundant empty lines to make the output text cleaner.
Fine-Tuning Tips
If you still see minor issues with specific PDFs, tweak the char_margin value (try 0.8–1.5). Every PDF has unique layout rules, so a small adjustment might be needed for edge cases.
内容的提问来源于stack exchange,提问作者Yue Zhao

