如何使用pdfminer获取字符字体样式、颜色及解析LTChar的matrix属性?
Hey there! Let me walk you through how to grab font styles, colors with pdfminer, and demystify that LTChar matrix property you're confused about.
First off, that matrix attribute is a PDF transformation matrix—think of it as the set of instructions that tells the renderer where to place the character, and how to scale/rotate/shear it. It’s a 6-element tuple (a, b, c, d, e, f) that maps to this 3x3 matrix (PDF uses homogeneous coordinates):
[ a b 0 ] [ c d 0 ] [ e f 1 ]
Here’s what each element does for a character:
a: Horizontal scaling factor (also interacts withbto affect vertical slant)b: Vertical shear factor (controls how much the character tilts vertically)c: Horizontal shear factor (controls horizontal slant)d: Vertical scaling factore: Horizontal translation (the X-coordinate of the character’s starting position)f: Vertical translation (the Y-coordinate of the character’s starting position)
For example, a standard untransformed character would have a matrix like (1, 0, 0, 1, 150, 300)—meaning no scaling/shearing, placed at (150, 300) on the page. If you see (0.9, 0, 0, 1.1, 200, 400), that character is squeezed 10% horizontally and stretched 10% vertically, positioned at (200,400).
To get font details from an LTChar, you’ll use its font property (an LTFont object). Here’s a practical code snippet:
from pdfminer.high_level import extract_pages from pdfminer.layout import LTChar for page in extract_pages("your_document.pdf"): for element in page: if isinstance(element, LTChar): # Get basic font info font_name = element.font.fontname font_size = element.size # This is the rendered size, accounting for matrix scaling # Check for bold/italic (two reliable methods) # Method 1: Check font name (works for most standard fonts) is_bold = "Bold" in font_name or "bold" in font_name is_italic = "Italic" in font_name or "italic" in font_name or "Oblique" in font_name # Method 2: Use font flags (more accurate for non-standard font names) # PDF font flags: 0x0002 = Bold, 0x0010 = Italic is_bold_flag = (element.font.flags & 0x0002) != 0 is_italic_flag = (element.font.flags & 0x0010) != 0 print(f"Font: {font_name}, Size: {font_size:.2f}") print(f"Bold (name check): {is_bold}, Bold (flag check): {is_bold_flag}") print(f"Italic (name check): {is_italic}, Italic (flag check): {is_italic_flag}")
The flag method is better for edge cases where font names don’t include "Bold" or "Italic" explicitly.
Text color in PDFs is controlled by fill and stroke states—most text uses fill color (solid text), while stroke color is for outlined text. Here’s how to access it:
for page in extract_pages("your_document.pdf"): for element in page: if isinstance(element, LTChar): # Fill color is the primary text color (returns None for default black) fill_color = element.fill_color # Stroke color is for text outlines (rarely used for regular text) stroke_color = element.stroke_color if fill_color: # Convert from 0-1 RGB range to 0-255 for readability rgb_color = tuple(int(channel * 255) for channel in fill_color) print(f"Text fill color (RGB): {rgb_color}") else: print("Text color: Default black") if stroke_color: rgb_stroke = tuple(int(channel * 255) for channel in stroke_color) print(f"Text stroke color (RGB): {rgb_stroke}")
Note: If an LTChar doesn’t have its own color set, it might inherit the color from its parent elements (like LTTextLine or LTTextBox). If you don’t find a color on the LTChar, you can traverse up the parent hierarchy to check.
内容的提问来源于stack exchange,提问作者Neha Pandey

