如何从多级列表标题PDF提取最低级标题及对应文本并导入Excel?
Approach Recommendation
Regex is a reliable choice for this task, especially since your headings follow a numbered hierarchical structure. It’s straightforward to match the pattern of lowest-level headings and capture their associated content. However, if the PDF conversion introduced inconsistent formatting (like line breaks in headings), a PDF-specific parser like pdfplumber might yield better results.
Regex-Based Solution (Using Converted Text File)
This script will parse your output.txt file, identify the deepest (lowest-level) numbered headings, extract their adjacent text, and save the pairs to a CSV file.
Step 1: Python Code
import re import csv # Read the converted text file with open('output.txt', 'r', encoding='utf-8') as f: text = f.read() # Regex pattern to match numbered headings (e.g., 1.2.3) and their subsequent content pattern = r'^(\d+(\.\d+)+)\s+(.+?)\n([\s\S]*?)(?=^\d+(\.\d+)+|$)' matches = re.findall(pattern, text, flags=re.MULTILINE) # Process matches to get heading title, content, and hierarchy depth entries = [] for match in matches: heading_num, _, heading_title, content, _ = match # Calculate depth (number of levels in the heading) depth = len(heading_num.split('.')) # Clean up content: remove extra newlines and whitespace cleaned_content = re.sub(r'\s+', ' ', content).strip() entries.append((depth, heading_title.strip(), cleaned_content)) if not entries: print("No numbered headings found in the text file.") exit() # Identify the lowest level (maximum depth) max_depth = max(entry[0] for entry in entries) # Filter only entries from the lowest level lowest_level_entries = [entry[1:] for entry in entries if entry[0] == max_depth] # Write to CSV (columns will map to Excel's C and D) with open('output.csv', 'w', newline='', encoding='utf-8') as csvfile: writer = csv.writer(csvfile) # Optional header (Excel will ignore this if you start importing at column C) writer.writerow(['Heading', 'Text']) for title, content in lowest_level_entries: writer.writerow([title, content]) print(f"Successfully exported {len(lowest_level_entries)} lowest-level entries to output.csv")
Key Details
- Regex Pattern: Matches headings with numeric hierarchy (e.g.,
1.2.3), captures the heading text, then grabs all content until the next heading or end of file. - Depth Calculation: Determines the lowest level by counting the number of dots in the heading number (more dots = deeper level).
- Content Cleaning: Normalizes whitespace to ensure clean text in Excel.
Alternative: PDF Parser Approach (No Pre-Conversion Needed)
If the text conversion introduced formatting issues, use pdfplumber to directly extract structured content from the PDF. This method relies on font size or indentation to identify headings.
Step 1: Install Dependencies
pip install pdfplumber
Step 2: Python Code
import pdfplumber import csv lowest_level_entries = [] # Adjust these values to match your PDF's formatting HEADING_FONT_SIZE = 12 # Font size of lowest-level headings CONTENT_FONT_SIZE = 10 # Font size of body text with pdfplumber.open('doc.pdf') as pdf: for page in pdf.pages: # Extract words with font size metadata words = page.extract_words(extra_attrs=['fontsize']) current_heading = None current_content = [] for word in words: # Check if the word belongs to a lowest-level heading if word['fontsize'] == HEADING_FONT_SIZE: # Save the previous heading-content pair if exists if current_heading is not None: cleaned_content = ' '.join(current_content).strip() lowest_level_entries.append((current_heading, cleaned_content)) current_heading = word['text'] current_content = [] # Add words to content if we're in a heading's section elif current_heading is not None and word['fontsize'] == CONTENT_FONT_SIZE: current_content.append(word['text']) # Save the last entry on the page if current_heading is not None: cleaned_content = ' '.join(current_content).strip() lowest_level_entries.append((current_heading, cleaned_content)) # Write to CSV with open('output.csv', 'w', newline='', encoding='utf-8') as csvfile: writer = csv.writer(csvfile) writer.writerow(['Heading', 'Text']) for title, content in lowest_level_entries: writer.writerow([title, content]) print(f"Successfully exported {len(lowest_level_entries)} entries to output.csv")
Key Details
- Font Size Detection: Adjust
HEADING_FONT_SIZEandCONTENT_FONT_SIZEto match your PDF’s actual formatting (usepdfplumberto inspect font sizes if unsure). - Structural Preservation: Directly reads the PDF, avoiding issues from text conversion like broken headings.
Excel Import Tips
When importing the CSV into Excel:
- Go to Data > From Text/CSV.
- Select
output.csvand click Import. - In the wizard, choose Delimited and set the delimiter to comma.
- Under Data preview, use the arrow buttons to skip columns A and B (so the first column maps to C, second to D).
- Click Load to complete the import.
内容的提问来源于stack exchange,提问作者wuddadid

