如何统计Word(.docx)文档自动换行后的标签对应行数?
Automated Line Counting for .docx with Text-Wrap Indentation
Great question—manual font and line length inputs are definitely a pain, especially when dealing with dynamic docx content. Here's a fully automated approach that extracts all necessary metadata directly from your .docx file and calculates line counts accurately, no manual parameters needed.
Core Approach
Instead of guessing font details or line lengths, we'll:
- Extract font properties (name, size, style) from the docx's underlying XML via
python-docx. - Retrieve page layout info (total width, margins) to determine the actual usable line width where text wraps.
- Calculate rendered text width for each text segment using PIL, matching the exact font from the docx.
- Compute line counts by dividing total text width by usable line width (rounding up for partial lines).
- Aggregate counts per tag based on your document's tag structure.
Step-by-Step Implementation
1. Install Required Libraries
First, install the tools we'll need:
pip install python-docx pillow python-fontconfig # Skip python-fontconfig on Windows
2. Extract Docx Metadata
Use python-docx to pull font and page info:
from docx import Document from docx.shared import Pt def get_docx_metadata(doc_path): doc = Document(doc_path) # Calculate usable line width (total page width minus left/right margins) section = doc.sections[0] total_width = section.page_width.pt left_margin = section.left_margin.pt right_margin = section.right_margin.pt usable_width = total_width - left_margin - right_margin # Collect font details for each paragraph font_details = [] for para in doc.paragraphs: if para.runs: run = para.runs[0] font_name = run.font.name or "Arial" # Fallback to default if missing font_size = run.font.size.pt if run.font.size else 12 font_bold = run.font.bold or False font_italic = run.font.italic or False font_details.append((font_name, font_size, font_bold, font_italic)) return usable_width, font_details, doc
3. Find System Font Path (Cross-Platform)
Map the docx's font name to the actual system font file for PIL:
import sys from PIL import ImageFont def get_font_path(font_name, bold=False, italic=False): # Windows font lookup if sys.platform == 'win32': import win32api fonts = win32api.GetFontData() for font in fonts: font_family = font[1].lower() if font_name.lower() in font_family: if bold and 'bold' in font_family: return font[0] if italic and 'italic' in font_family: return font[0] if not bold and not italic: return font[0] return ImageFont.truetype("arial.ttf", 12) # Linux/macOS using fontconfig else: import fontconfig pattern = fontconfig.FcPattern() pattern.add("family", font_name) if bold: pattern.add("weight", fontconfig.FcWeight.BOLD) if italic: pattern.add("slant", fontconfig.FcSlant.ITALIC) fonts = fontconfig.FcFontList(None, pattern, None) return fonts[0]['file'] if fonts else ImageFont.truetype("DejaVuSans.ttf", 12)
4. Calculate Line Count for Text
Use PIL to measure text width and compute line breaks:
from math import ceil def calculate_line_count(text, font_path, font_size, usable_width): font = ImageFont.truetype(font_path, font_size) text_width = font.getlength(text) # Round up to count partial lines as full lines return ceil(text_width / usable_width)
5. Aggregate Counts per Tag
Assuming your document uses tags like [TagName] to mark sections, here's how to split and count:
def count_lines_per_tag(doc_path): usable_width, font_details, doc = get_docx_metadata(doc_path) tag_counts = {} current_tag = None current_text = "" for idx, para in enumerate(doc.paragraphs): # Detect new tag (adjust regex to match your tag format) para_text = para.text.strip() if para_text.startswith('[') and para_text.endswith(']'): # Save previous tag's count if current_tag and current_text: font_name, size, bold, italic = font_details[idx-1] font_path = get_font_path(font_name, bold, italic) line_count = calculate_line_count(current_text, font_path, size, usable_width) tag_counts[current_tag] = tag_counts.get(current_tag, 0) + line_count # Set new tag current_tag = para_text[1:-1] current_text = "" elif current_tag: current_text += para.text + " " # Preserve spacing between paragraphs # Handle the last tag in the document if current_tag and current_text: font_name, size, bold, italic = font_details[-1] font_path = get_font_path(font_name, bold, italic) line_count = calculate_line_count(current_text, font_path, size, usable_width) tag_counts[current_tag] = tag_counts.get(current_tag, 0) + line_count return tag_counts
6. Run the Script
if __name__ == "__main__": doc_path = "your_document.docx" results = count_lines_per_tag(doc_path) for tag, count in results.items(): print(f"Tag {tag}: {count} lines")
Notes & Edge Cases
- Multiple Fonts in a Tag: If a tag uses mixed fonts/sizes, modify the script to calculate line counts per text run instead of per paragraph.
- Custom Tag Formats: Update the tag detection logic to match your actual tag syntax (e.g.,
### TagName ###instead of[TagName]). - Fallback Fonts: The script includes fallback fonts if the docx's font isn't found on your system—adjust these to match your setup.
内容的提问来源于stack exchange,提问作者Daniele Grattarola
相关产品推荐
相关产品推荐

