You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从多级列表标题PDF提取最低级标题及对应文本并导入Excel?

Extracting Lowest-Level Headings and Content from PDF Text

Approach Recommendation

Regex is a reliable choice for this task, especially since your headings follow a numbered hierarchical structure. It’s straightforward to match the pattern of lowest-level headings and capture their associated content. However, if the PDF conversion introduced inconsistent formatting (like line breaks in headings), a PDF-specific parser like pdfplumber might yield better results.

Regex-Based Solution (Using Converted Text File)

This script will parse your output.txt file, identify the deepest (lowest-level) numbered headings, extract their adjacent text, and save the pairs to a CSV file.

Step 1: Python Code

import re
import csv

# Read the converted text file
with open('output.txt', 'r', encoding='utf-8') as f:
    text = f.read()

# Regex pattern to match numbered headings (e.g., 1.2.3) and their subsequent content
pattern = r'^(\d+(\.\d+)+)\s+(.+?)\n([\s\S]*?)(?=^\d+(\.\d+)+|$)'
matches = re.findall(pattern, text, flags=re.MULTILINE)

# Process matches to get heading title, content, and hierarchy depth
entries = []
for match in matches:
    heading_num, _, heading_title, content, _ = match
    # Calculate depth (number of levels in the heading)
    depth = len(heading_num.split('.'))
    # Clean up content: remove extra newlines and whitespace
    cleaned_content = re.sub(r'\s+', ' ', content).strip()
    entries.append((depth, heading_title.strip(), cleaned_content))

if not entries:
    print("No numbered headings found in the text file.")
    exit()

# Identify the lowest level (maximum depth)
max_depth = max(entry[0] for entry in entries)

# Filter only entries from the lowest level
lowest_level_entries = [entry[1:] for entry in entries if entry[0] == max_depth]

# Write to CSV (columns will map to Excel's C and D)
with open('output.csv', 'w', newline='', encoding='utf-8') as csvfile:
    writer = csv.writer(csvfile)
    # Optional header (Excel will ignore this if you start importing at column C)
    writer.writerow(['Heading', 'Text'])
    for title, content in lowest_level_entries:
        writer.writerow([title, content])

print(f"Successfully exported {len(lowest_level_entries)} lowest-level entries to output.csv")

Key Details

  • Regex Pattern: Matches headings with numeric hierarchy (e.g., 1.2.3), captures the heading text, then grabs all content until the next heading or end of file.
  • Depth Calculation: Determines the lowest level by counting the number of dots in the heading number (more dots = deeper level).
  • Content Cleaning: Normalizes whitespace to ensure clean text in Excel.

Alternative: PDF Parser Approach (No Pre-Conversion Needed)

If the text conversion introduced formatting issues, use pdfplumber to directly extract structured content from the PDF. This method relies on font size or indentation to identify headings.

Step 1: Install Dependencies

pip install pdfplumber

Step 2: Python Code

import pdfplumber
import csv

lowest_level_entries = []

# Adjust these values to match your PDF's formatting
HEADING_FONT_SIZE = 12  # Font size of lowest-level headings
CONTENT_FONT_SIZE = 10  # Font size of body text

with pdfplumber.open('doc.pdf') as pdf:
    for page in pdf.pages:
        # Extract words with font size metadata
        words = page.extract_words(extra_attrs=['fontsize'])
        current_heading = None
        current_content = []
        
        for word in words:
            # Check if the word belongs to a lowest-level heading
            if word['fontsize'] == HEADING_FONT_SIZE:
                # Save the previous heading-content pair if exists
                if current_heading is not None:
                    cleaned_content = ' '.join(current_content).strip()
                    lowest_level_entries.append((current_heading, cleaned_content))
                current_heading = word['text']
                current_content = []
            # Add words to content if we're in a heading's section
            elif current_heading is not None and word['fontsize'] == CONTENT_FONT_SIZE:
                current_content.append(word['text'])
        
        # Save the last entry on the page
        if current_heading is not None:
            cleaned_content = ' '.join(current_content).strip()
            lowest_level_entries.append((current_heading, cleaned_content))

# Write to CSV
with open('output.csv', 'w', newline='', encoding='utf-8') as csvfile:
    writer = csv.writer(csvfile)
    writer.writerow(['Heading', 'Text'])
    for title, content in lowest_level_entries:
        writer.writerow([title, content])

print(f"Successfully exported {len(lowest_level_entries)} entries to output.csv")

Key Details

  • Font Size Detection: Adjust HEADING_FONT_SIZE and CONTENT_FONT_SIZE to match your PDF’s actual formatting (use pdfplumber to inspect font sizes if unsure).
  • Structural Preservation: Directly reads the PDF, avoiding issues from text conversion like broken headings.

Excel Import Tips

When importing the CSV into Excel:

  • Go to Data > From Text/CSV.
  • Select output.csv and click Import.
  • In the wizard, choose Delimited and set the delimiter to comma.
  • Under Data preview, use the arrow buttons to skip columns A and B (so the first column maps to C, second to D).
  • Click Load to complete the import.

内容的提问来源于stack exchange,提问作者wuddadid

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:14:30