Python提取多格式分隔符间文本及批量处理的技术求助
Hey there! As someone with R experience, you’re already familiar with pattern matching—let’s adapt that to Python to solve your large-scale text extraction problem. Here’s how to address all three of your questions:
1. Generalized Pattern Matching for All Variants
The key here is to build a regex flexible enough to handle every Item 7/Item 7A variation you’ve encountered (case differences, punctuation, line breaks). We’ll use:
re.IGNORECASEto match bothItemandITEMre.DOTALLto let.match newline characters (so multi-line content gets captured)- A pattern that accounts for optional punctuation (
.or:) after7and7A, plus any amount of whitespace
The regex pattern looks like this:
import re pattern = re.compile(r'ITEM 7(?:\.|:)?\s*(.*?)ITEM 7A(?:\.|:)?', re.IGNORECASE | re.DOTALL)
Quick breakdown:
ITEM 7: Matches "Item 7" or "ITEM 7" (case-insensitive)(?:\.|:)?: Optional non-capturing group for.or:after the 7\s*: Matches any whitespace (including newlines) between the header and content(.*?): Non-greedy capture of all content until we hit the closing headerITEM 7A(?:\.|:)?: Matches the closing header with optional punctuation
2. Save Each Extracted Block to Separate Files
For every input .txt file, we’ll extract the matching content and save it to a unique output file (e.g., if your input is report_123.txt, the output becomes report_123_item7_content.txt). If a file has multiple matches (unlikely but possible), we’ll append a number to avoid overwriting.
3. Generate a Log for Unprocessed Files
We’ll track two types of issues:
- Files where no
Item 7/Item 7Ablock was found - Files that threw errors during reading (e.g., permission issues, corrupted data)
The log will include timestamps and clear error messages to help you follow up on problematic files.
Complete Working Code
import glob import os import re from datetime import datetime def setup_paths(input_dir, output_dir, log_path): # Create output directory if it doesn't exist os.makedirs(output_dir, exist_ok=True) # Initialize log file with header with open(log_path, 'w', encoding='utf-8') as log_file: log_file.write(f"Log started at: {datetime.now().strftime('%Y-%m-%d %H:%M:%S')}\n") log_file.write("="*50 + "\n") def extract_and_save(input_file, output_dir, pattern, log_file): try: with open(input_file, 'r', encoding='utf-8') as f: data = f.read() matches = pattern.findall(data) if not matches: log_file.write(f"No match found: {input_file}\n") return # Get base filename without extension base_name = os.path.splitext(os.path.basename(input_file))[0] for idx, content in enumerate(matches, 1): # Clean up content: replace newlines with spaces, strip extra whitespace cleaned_content = content.replace('\n', ' ').strip() # Define output filename output_filename = f"{base_name}_item7_content{'_' + str(idx) if len(matches) > 1 else ''}.txt" output_path = os.path.join(output_dir, output_filename) # Write cleaned content to file with open(output_path, 'w', encoding='utf-8') as out_file: out_file.write(cleaned_content) except Exception as e: log_file.write(f"Error processing {input_file}: {str(e)}\n") def main(): # Configure your paths here input_dir = r"C:\R_Practice\Your_Input_Files" # Replace with your input folder output_dir = r"C:\R_Practice\Extracted_Item7" # Replace with your output folder log_path = r"C:\R_Practice\extraction_log.txt" # Replace with your log file path # Setup initial paths and log setup_paths(input_dir, output_dir, log_path) # Compile regex once (efficient for large batches) pattern = re.compile(r'ITEM 7(?:\.|:)?\s*(.*?)ITEM 7A(?:\.|:)?', re.IGNORECASE | re.DOTALL) # Process all .txt files with open(log_path, 'a', encoding='utf-8') as log_file: for filename in glob.glob(os.path.join(input_dir, '*.txt')): extract_and_save(filename, output_dir, pattern, log_file) log_file.write("\n" + "="*50 + "\n") log_file.write(f"Log finished at: {datetime.now().strftime('%Y-%m-%d %H:%M:%S')}\n") if __name__ == "__main__": main()
Key Tips for Your 250k File Batch:
- Efficiency: Compiling the regex once (outside the loop) saves significant time when processing thousands of files.
- Encoding: Using
encoding='utf-8'avoids common Windows text file encoding errors. - Scalability: Make sure your output directory has enough disk space—250k small text files add up!
内容的提问来源于stack exchange,提问作者Cerberus_64

