pdftotext转换PDF出现单词粘连的原因及解决办法
Hey there, I’ve dealt with this exact word-concatenation issue when extracting text from PDFs—total pain when it ruins the readability of your output. Let’s walk through why this happens and the fixes I’ve found work.
Why This Happens
- PDF Layout vs. Actual Text Structure: A lot of PDFs don’t store text as plain, spaced words. Instead, they position text blocks using x/y coordinates. If the spacing between words is implemented as a tiny coordinate shift (not an actual space character), tools like
pdftotextor Okular might misinterpret that shift as no space at all, gluing words together. - Font or Kerning Issues: If the PDF uses non-standard embedded fonts, or has aggressive kerning (tight spacing between characters), extraction tools can’t distinguish between intentional tight letter spacing and missing word spaces.
- Scanned/OCR’d PDFs: If your PDF is a scanned image that was OCR’d later, the OCR engine might have failed to detect spaces between words—especially with small fonts, blurry text, or complex layouts.
Fixes to Try
1. Adjust pdftotext Parameters
The default pdftotext mode prioritizes flow over layout, which can break spacing. Try these flags:
- Use the
-layoutflag to preserve the original PDF’s layout structure. This is usually the first fix to test:pdftotext -layout your_file.pdf output.txt - Use
-fixed Nto force a fixed character width (replaceNwith a value between 6-10, depending on your PDF’s font size). This helps override wonky kerning:pdftotext -fixed 8 your_file.pdf output.txt
2. Switch to a Different Extraction Tool
If pdftotext isn’t cutting it, try these more robust alternatives:
- pdfminer.six’s
pdf2txt.py: It has better layout analysis for tricky PDFs. Install it viapip install pdfminer.six, then run:pdf2txt.py -o output.txt your_file.pdf - MuPDF’s
mutool: This tool’s text extraction handles spacing issues well for many cases. Run:mutool convert -o output.txt your_file.pdf
3. Post-Extraction Text Repair
If you already have the messed-up text, use regex to fix common粘连 cases (test these carefully to avoid false positives):
- Fix letter-number concatenation (e.g.,
the2nd→the 2nd):import re corrected_text = re.sub(r'([a-zA-Z])(\d)', r'\1 \2', your_text) - For all-lowercase word粘连 (like
beforeme), you’ll need a more cautious approach—maybe use a spell checker to identify valid words and split where needed, or use a tool liketextblobto detect word boundaries.
4. Re-OCR Scanned PDFs
If your PDF is a scanned image, re-run OCR with a more accurate engine like Tesseract:
tesseract your_file.pdf output.txt --oem 3 --psm 6
The --oem 3 uses the default OCR engine mode, and --psm 6 tells Tesseract to treat the input as a single uniform block of text, which can improve space detection.
内容的提问来源于stack exchange,提问作者anushka

