You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python提取多格式分隔符间文本及批量处理的技术求助

Hey there! As someone with R experience, you’re already familiar with pattern matching—let’s adapt that to Python to solve your large-scale text extraction problem. Here’s how to address all three of your questions:

1. Generalized Pattern Matching for All Variants

The key here is to build a regex flexible enough to handle every Item 7/Item 7A variation you’ve encountered (case differences, punctuation, line breaks). We’ll use:

  • re.IGNORECASE to match both Item and ITEM
  • re.DOTALL to let . match newline characters (so multi-line content gets captured)
  • A pattern that accounts for optional punctuation (. or :) after 7 and 7A, plus any amount of whitespace

The regex pattern looks like this:

import re
pattern = re.compile(r'ITEM 7(?:\.|:)?\s*(.*?)ITEM 7A(?:\.|:)?', re.IGNORECASE | re.DOTALL)

Quick breakdown:

  • ITEM 7: Matches "Item 7" or "ITEM 7" (case-insensitive)
  • (?:\.|:)?: Optional non-capturing group for . or : after the 7
  • \s*: Matches any whitespace (including newlines) between the header and content
  • (.*?): Non-greedy capture of all content until we hit the closing header
  • ITEM 7A(?:\.|:)?: Matches the closing header with optional punctuation

2. Save Each Extracted Block to Separate Files

For every input .txt file, we’ll extract the matching content and save it to a unique output file (e.g., if your input is report_123.txt, the output becomes report_123_item7_content.txt). If a file has multiple matches (unlikely but possible), we’ll append a number to avoid overwriting.

3. Generate a Log for Unprocessed Files

We’ll track two types of issues:

  • Files where no Item 7/Item 7A block was found
  • Files that threw errors during reading (e.g., permission issues, corrupted data)

The log will include timestamps and clear error messages to help you follow up on problematic files.


Complete Working Code

import glob
import os
import re
from datetime import datetime

def setup_paths(input_dir, output_dir, log_path):
    # Create output directory if it doesn't exist
    os.makedirs(output_dir, exist_ok=True)
    # Initialize log file with header
    with open(log_path, 'w', encoding='utf-8') as log_file:
        log_file.write(f"Log started at: {datetime.now().strftime('%Y-%m-%d %H:%M:%S')}\n")
        log_file.write("="*50 + "\n")

def extract_and_save(input_file, output_dir, pattern, log_file):
    try:
        with open(input_file, 'r', encoding='utf-8') as f:
            data = f.read()
        
        matches = pattern.findall(data)
        if not matches:
            log_file.write(f"No match found: {input_file}\n")
            return
        
        # Get base filename without extension
        base_name = os.path.splitext(os.path.basename(input_file))[0]
        
        for idx, content in enumerate(matches, 1):
            # Clean up content: replace newlines with spaces, strip extra whitespace
            cleaned_content = content.replace('\n', ' ').strip()
            # Define output filename
            output_filename = f"{base_name}_item7_content{'_' + str(idx) if len(matches) > 1 else ''}.txt"
            output_path = os.path.join(output_dir, output_filename)
            
            # Write cleaned content to file
            with open(output_path, 'w', encoding='utf-8') as out_file:
                out_file.write(cleaned_content)
                
    except Exception as e:
        log_file.write(f"Error processing {input_file}: {str(e)}\n")

def main():
    # Configure your paths here
    input_dir = r"C:\R_Practice\Your_Input_Files"  # Replace with your input folder
    output_dir = r"C:\R_Practice\Extracted_Item7"   # Replace with your output folder
    log_path = r"C:\R_Practice\extraction_log.txt"  # Replace with your log file path
    
    # Setup initial paths and log
    setup_paths(input_dir, output_dir, log_path)
    
    # Compile regex once (efficient for large batches)
    pattern = re.compile(r'ITEM 7(?:\.|:)?\s*(.*?)ITEM 7A(?:\.|:)?', re.IGNORECASE | re.DOTALL)
    
    # Process all .txt files
    with open(log_path, 'a', encoding='utf-8') as log_file:
        for filename in glob.glob(os.path.join(input_dir, '*.txt')):
            extract_and_save(filename, output_dir, pattern, log_file)
        
        log_file.write("\n" + "="*50 + "\n")
        log_file.write(f"Log finished at: {datetime.now().strftime('%Y-%m-%d %H:%M:%S')}\n")

if __name__ == "__main__":
    main()

Key Tips for Your 250k File Batch:

  • Efficiency: Compiling the regex once (outside the loop) saves significant time when processing thousands of files.
  • Encoding: Using encoding='utf-8' avoids common Windows text file encoding errors.
  • Scalability: Make sure your output directory has enough disk space—250k small text files add up!

内容的提问来源于stack exchange,提问作者Cerberus_64

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:51:26