Python多文件正则匹配需求及问题求助(正则熟但Python新手)
Hey there! Great job diving into Python with your regex expertise—6 hours is a solid start. Let's build a script that checks off all your requirements and fixes those current issues you mentioned. Here's a step-by-step solution tailored to your needs:
Core Features We'll Implement
- Traverse all text files in a folder (easy to extend to docx/xml later)
- Load regex patterns from an external text file (or use a single pattern)
- Capture all matching details (filename, full match, capture groups, line content)
- Split multiple matches per line into separate rows
- Export results to a Pandas DataFrame or CSV for further analysis
Full Solution Code
import re import os import pandas as pd def load_regex_patterns(pattern_file_path): """Load regex patterns from a text file (one pattern per line)""" patterns = [] with open(pattern_file_path, 'r', encoding='utf-8') as f: for line in f: pattern = line.strip() if pattern: # Skip empty lines patterns.append(re.compile(pattern)) return patterns def process_file(file_path, regex_patterns): """Process a single file, collect all matches with details""" matches_data = [] with open(file_path, 'r', encoding='utf-8') as f: filename = os.path.basename(file_path) for line_num, line_content in enumerate(f, start=1): line_content = line_content.strip() if not line_content: continue # Skip empty lines for regex in regex_patterns: # Find all matches (including capture groups) all_matches = regex.finditer(line_content) for match in all_matches: # Extract full match and capture groups full_match = match.group(0) capture_groups = match.groups() # Tuple of captured groups # Add to our data list as a dictionary matches_data.append({ 'filename': filename, 'line_number': line_num, 'line_content': line_content, 'full_match': full_match, 'capture_groups': capture_groups # Store as tuple, can convert to string if needed }) return matches_data def main(folder_path, pattern_file_path=None, single_pattern=None): """Main function to orchestrate the process""" # Load regex patterns if pattern_file_path: regex_patterns = load_regex_patterns(pattern_file_path) elif single_pattern: regex_patterns = [re.compile(single_pattern)] else: raise ValueError("Either provide a pattern file path or a single regex pattern") # Traverse all .txt files in the folder (and subfolders) all_matches = [] for root, dirs, files in os.walk(folder_path): for file in files: if file.lower().endswith('.txt'): file_path = os.path.join(root, file) print(f"Processing file: {file_path}") file_matches = process_file(file_path, regex_patterns) all_matches.extend(file_matches) # Convert to Pandas DataFrame for analysis df = pd.DataFrame(all_matches) # Optional: Save to CSV df.to_csv('regex_matches_results.csv', index=False, encoding='utf-8') print(f"Processing complete! Total matches found: {len(all_matches)}") return df # Example usage if __name__ == "__main__": # Use either a single pattern or a pattern file # Option 1: Single pattern # result_df = main(folder_path='your_target_folder', single_pattern=r"Bob(by)?") # Option 2: Load from pattern file (one regex per line) result_df = main(folder_path='your_target_folder', pattern_file_path='regex_patterns.txt') # You can now use result_df for further analysis in Pandas print(result_df.head())
Key Fixes & Improvements Explained
Let's break down how this solves your specific issues:
1. Working with Capture Groups
- We use
regex.finditer()instead offindall()because it returns match objects, which let us access both the full match (match.group(0)) and individual capture groups (match.groups()). - Capture groups are stored as a tuple in the DataFrame—you can easily convert them to strings or split into separate columns if needed (e.g.,
df[['group1', 'group2']] = pd.DataFrame(df['capture_groups'].tolist())).
2. Reusing Data in Pandas
- All match details are collected into a list of dictionaries, which converts seamlessly into a Pandas DataFrame.
- Once you have
result_df, you can filter, sort, or analyze the data however you want—for example, count matches per filename, or find which lines have the most matches.
3. Split Multiple Matches into Separate Rows
- By looping over each match from
finditer(), every match in a single line gets its own entry in the data list. For your example withBob(by)?matching "We used to call Bob 'Little Bobby'", you'll get two rows: one for "Bob" and one for "Bobby", each linked to the same line content and filename.
Next Steps for Extending to Docx/XML
Since you mentioned handling docx/xml later, here's a quick tip:
- For
.docxfiles, use thepython-docxlibrary to extract text from documents. - For
.xmlfiles, use Python's built-inxml.etree.ElementTreeorlxmlto parse and extract text content. - You can modify the
process_filefunction to detect file extensions and use the appropriate text extraction method.
内容的提问来源于stack exchange,提问作者CaseLawMiner
相关产品推荐
相关产品推荐

