You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python多文件正则匹配需求及问题求助(正则熟但Python新手)

Hey there! Great job diving into Python with your regex expertise—6 hours is a solid start. Let's build a script that checks off all your requirements and fixes those current issues you mentioned. Here's a step-by-step solution tailored to your needs:

Core Features We'll Implement

  • Traverse all text files in a folder (easy to extend to docx/xml later)
  • Load regex patterns from an external text file (or use a single pattern)
  • Capture all matching details (filename, full match, capture groups, line content)
  • Split multiple matches per line into separate rows
  • Export results to a Pandas DataFrame or CSV for further analysis

Full Solution Code

import re
import os
import pandas as pd

def load_regex_patterns(pattern_file_path):
    """Load regex patterns from a text file (one pattern per line)"""
    patterns = []
    with open(pattern_file_path, 'r', encoding='utf-8') as f:
        for line in f:
            pattern = line.strip()
            if pattern:  # Skip empty lines
                patterns.append(re.compile(pattern))
    return patterns

def process_file(file_path, regex_patterns):
    """Process a single file, collect all matches with details"""
    matches_data = []
    with open(file_path, 'r', encoding='utf-8') as f:
        filename = os.path.basename(file_path)
        for line_num, line_content in enumerate(f, start=1):
            line_content = line_content.strip()
            if not line_content:
                continue  # Skip empty lines
            
            for regex in regex_patterns:
                # Find all matches (including capture groups)
                all_matches = regex.finditer(line_content)
                for match in all_matches:
                    # Extract full match and capture groups
                    full_match = match.group(0)
                    capture_groups = match.groups()  # Tuple of captured groups
                    
                    # Add to our data list as a dictionary
                    matches_data.append({
                        'filename': filename,
                        'line_number': line_num,
                        'line_content': line_content,
                        'full_match': full_match,
                        'capture_groups': capture_groups  # Store as tuple, can convert to string if needed
                    })
    return matches_data

def main(folder_path, pattern_file_path=None, single_pattern=None):
    """Main function to orchestrate the process"""
    # Load regex patterns
    if pattern_file_path:
        regex_patterns = load_regex_patterns(pattern_file_path)
    elif single_pattern:
        regex_patterns = [re.compile(single_pattern)]
    else:
        raise ValueError("Either provide a pattern file path or a single regex pattern")
    
    # Traverse all .txt files in the folder (and subfolders)
    all_matches = []
    for root, dirs, files in os.walk(folder_path):
        for file in files:
            if file.lower().endswith('.txt'):
                file_path = os.path.join(root, file)
                print(f"Processing file: {file_path}")
                file_matches = process_file(file_path, regex_patterns)
                all_matches.extend(file_matches)
    
    # Convert to Pandas DataFrame for analysis
    df = pd.DataFrame(all_matches)
    
    # Optional: Save to CSV
    df.to_csv('regex_matches_results.csv', index=False, encoding='utf-8')
    
    print(f"Processing complete! Total matches found: {len(all_matches)}")
    return df

# Example usage
if __name__ == "__main__":
    # Use either a single pattern or a pattern file
    # Option 1: Single pattern
    # result_df = main(folder_path='your_target_folder', single_pattern=r"Bob(by)?")
    
    # Option 2: Load from pattern file (one regex per line)
    result_df = main(folder_path='your_target_folder', pattern_file_path='regex_patterns.txt')
    
    # You can now use result_df for further analysis in Pandas
    print(result_df.head())

Key Fixes & Improvements Explained

Let's break down how this solves your specific issues:

1. Working with Capture Groups

  • We use regex.finditer() instead of findall() because it returns match objects, which let us access both the full match (match.group(0)) and individual capture groups (match.groups()).
  • Capture groups are stored as a tuple in the DataFrame—you can easily convert them to strings or split into separate columns if needed (e.g., df[['group1', 'group2']] = pd.DataFrame(df['capture_groups'].tolist())).

2. Reusing Data in Pandas

  • All match details are collected into a list of dictionaries, which converts seamlessly into a Pandas DataFrame.
  • Once you have result_df, you can filter, sort, or analyze the data however you want—for example, count matches per filename, or find which lines have the most matches.

3. Split Multiple Matches into Separate Rows

  • By looping over each match from finditer(), every match in a single line gets its own entry in the data list. For your example with Bob(by)? matching "We used to call Bob 'Little Bobby'", you'll get two rows: one for "Bob" and one for "Bobby", each linked to the same line content and filename.

Next Steps for Extending to Docx/XML

Since you mentioned handling docx/xml later, here's a quick tip:

  • For .docx files, use the python-docx library to extract text from documents.
  • For .xml files, use Python's built-in xml.etree.ElementTree or lxml to parse and extract text content.
  • You can modify the process_file function to detect file extensions and use the appropriate text extraction method.

内容的提问来源于stack exchange,提问作者CaseLawMiner

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:31:16