求助:如何用PDFMiner.six提取PDF单行文本的字体信息?
Got it, let's work through this! Since you already have basic PDF-to-text extraction up and running, grabbing line-specific font info just means leaning into PDFMiner.six's layout analysis tools—instead of only pulling raw text, we'll interact with the structural objects that hold styling metadata.
How It Works
PDFMiner breaks down PDFs into a hierarchy of layout objects, and the ones we care about here are:
LTTextLine: Represents a single line of textLTChar: Individual character objects within a line, each carrying font name, size, and other style details
Every line is made up of these LTChar instances, so we can pull font data directly from them.
Step-by-Step Code Example
First, make sure you're using a pdfminer.six version compatible with Python 3.6.3—install it with:
pip install pdfminer.six==20201018
Then use this code to extract line-level font information:
from pdfminer.high_level import extract_pages from pdfminer.layout import LTTextContainer, LTTextLine, LTChar from collections import defaultdict def extract_line_fonts(pdf_path): for page_num, page_layout in enumerate(extract_pages(pdf_path), start=1): print(f"--- Page {page_num} ---") for element in page_layout: # Only process elements that contain text if isinstance(element, LTTextContainer): for text_line in element: if isinstance(text_line, LTTextLine): line_text = text_line.get_text().strip() if not line_text: continue # Track font usage in the line (handles mixed fonts) font_stats = defaultdict(int) for char in text_line: if isinstance(char, LTChar): # Use (font name, size) as a unique identifier font_key = (char.fontname, round(char.size, 2)) font_stats[font_key] += 1 # Get the most common font in the line if font_stats: most_common_font = max(font_stats, key=font_stats.get) font_name, font_size = most_common_font print(f"Line: {line_text}") print(f"Font: {font_name}, Size: {font_size}\n") # Run the function on your target PDF extract_line_fonts("your_document.pdf")
Key Notes
- Handling Mixed Fonts: The code tracks all fonts in a line and picks the most common one. If you need to account for multiple fonts in a single line (e.g., bold headings within a regular text line), you can modify it to group characters by their font properties and report each one.
- Precision: Font sizes are rounded to 2 decimal places for readability—adjust the rounding value if you need more precision.
- High-Level API: We use
extract_pages()because it simplifies traversing the PDF layout compared to the low-level parser/interpreter setup you might have used for basic text extraction.
This should give you the line-specific font name and size you're looking for. Let me know if you run into edge cases with your specific PDFs!
内容的提问来源于stack exchange,提问作者mrwadepro

