You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求助:如何用PDFMiner.six提取PDF单行文本的字体信息?

Extracting Font Details (Name & Size) for Individual Text Lines with PDFMiner.six

Got it, let's work through this! Since you already have basic PDF-to-text extraction up and running, grabbing line-specific font info just means leaning into PDFMiner.six's layout analysis tools—instead of only pulling raw text, we'll interact with the structural objects that hold styling metadata.

How It Works

PDFMiner breaks down PDFs into a hierarchy of layout objects, and the ones we care about here are:

  • LTTextLine: Represents a single line of text
  • LTChar: Individual character objects within a line, each carrying font name, size, and other style details

Every line is made up of these LTChar instances, so we can pull font data directly from them.

Step-by-Step Code Example

First, make sure you're using a pdfminer.six version compatible with Python 3.6.3—install it with:

pip install pdfminer.six==20201018

Then use this code to extract line-level font information:

from pdfminer.high_level import extract_pages
from pdfminer.layout import LTTextContainer, LTTextLine, LTChar
from collections import defaultdict

def extract_line_fonts(pdf_path):
    for page_num, page_layout in enumerate(extract_pages(pdf_path), start=1):
        print(f"--- Page {page_num} ---")
        for element in page_layout:
            # Only process elements that contain text
            if isinstance(element, LTTextContainer):
                for text_line in element:
                    if isinstance(text_line, LTTextLine):
                        line_text = text_line.get_text().strip()
                        if not line_text:
                            continue
                        
                        # Track font usage in the line (handles mixed fonts)
                        font_stats = defaultdict(int)
                        for char in text_line:
                            if isinstance(char, LTChar):
                                # Use (font name, size) as a unique identifier
                                font_key = (char.fontname, round(char.size, 2))
                                font_stats[font_key] += 1
                        
                        # Get the most common font in the line
                        if font_stats:
                            most_common_font = max(font_stats, key=font_stats.get)
                            font_name, font_size = most_common_font
                            print(f"Line: {line_text}")
                            print(f"Font: {font_name}, Size: {font_size}\n")

# Run the function on your target PDF
extract_line_fonts("your_document.pdf")

Key Notes

  • Handling Mixed Fonts: The code tracks all fonts in a line and picks the most common one. If you need to account for multiple fonts in a single line (e.g., bold headings within a regular text line), you can modify it to group characters by their font properties and report each one.
  • Precision: Font sizes are rounded to 2 decimal places for readability—adjust the rounding value if you need more precision.
  • High-Level API: We use extract_pages() because it simplifies traversing the PDF layout compared to the low-level parser/interpreter setup you might have used for basic text extraction.

This should give you the line-specific font name and size you're looking for. Let me know if you run into edge cases with your specific PDFs!

内容的提问来源于stack exchange,提问作者mrwadepro

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 12:17:15