You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用PDFMiner转换PDF时部分文档无单词间距问题求助

Fixing Missing Spaces in PDF-to-TXT Conversion with PDFMiner

Hey there! I’ve dealt with this exact frustrating issue before—those run-together words happen because some PDFs (like the academic paper you linked) don’t use actual space characters to separate words. Instead, they rely on adjusting character positions to create visual spacing, which PDFMiner’s default settings miss. Here’s how to fix it:

Why This Happens

Older or layout-focused PDFs (common in academic publishing) often skip literal spaces and use horizontal gaps between characters to mimic word separation. PDFMiner’s default LAParams (Layout Analysis Parameters) are tuned for standard PDFs, so it doesn’t recognize these gaps as spaces.

Modified Conversion Code

We’ll tweak the LAParams settings to tell PDFMiner to detect those visual gaps as word separators. Here’s the updated function, plus your existing download code:

from pdfminer.pdfinterp import PDFResourceManager, PDFPageInterpreter
from pdfminer.converter import TextConverter
from pdfminer.layout import LAParams
from pdfminer.pdfpage import PDFPage
from io import StringIO
import requests

def convert_pdf_to_txt(path):
    rsrcmgr = PDFResourceManager()
    retstr = StringIO()
    codec = 'utf-8'
    
    # Key fix: Adjust layout analysis parameters to detect visual word gaps
    laparams = LAParams(
        char_margin=1.0,  # Lower threshold: gaps larger than this = new word
        word_margin=0.5,  # Ensure small gaps between words aren't merged
        detect_vertical=True,  # Handle any vertical text edge cases
        line_overlap=0.5
    )
    
    device = TextConverter(rsrcmgr, retstr, codec=codec, laparams=laparams)
    fp = open(path, 'rb')
    interpreter = PDFPageInterpreter(rsrcmgr, device)
    
    password = ""
    maxpages = 0
    caching = True
    pagenos = set()
    
    for page in PDFPage.get_pages(fp, pagenos, maxpages=maxpages, 
                                  password=password, caching=caching, 
                                  check_extractable=True):
        interpreter.process_page(page)
    
    fp.close()
    device.close()
    text = retstr.getvalue()
    retstr.close()
    
    # Optional: Clean up messy newlines and extra spaces
    text = '\n'.join([line.strip() for line in text.split('\n') if line.strip()])
    return text

# Your working PDF download code
url = 'http://www.ece.rochester.edu/~gsharma/papers/LocalImageRegisterEI2005.pdf'
file_name = './LocalImageRegisterEI2005.pdf'
response = requests.get(url)
with open(file_name, 'wb') as f:
    f.write(response.content)

# Test the conversion
converted_text = convert_pdf_to_txt(file_name)
print(converted_text[:500])  # Print first 500 chars to verify

What Each Parameter Does

  • char_margin: Default is 2.0. By lowering it to 1.0, we tell PDFMiner that any gap between characters larger than 1.0 units should be treated as a word separator (inserting a space). This is the main fix for your issue.
  • word_margin: Adjusts how close two words can be before being merged into one. 0.5 is a safe middle ground for most academic PDFs.
  • detect_vertical: Ensures we don’t miss text laid out vertically, which can cause weird merging in some documents.
  • The optional cleanup step removes redundant empty lines to make the output text cleaner.

Fine-Tuning Tips

If you still see minor issues with specific PDFs, tweak the char_margin value (try 0.8–1.5). Every PDF has unique layout rules, so a small adjustment might be needed for edge cases.

内容的提问来源于stack exchange,提问作者Yue Zhao

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 04:02:01