Python:在TXT文件中查找并统计单词的精确匹配与近似匹配
Solution for Approximate Keyword Matching in Text Files
Absolutely, this is totally achievable! The core fix here is swapping out exact string counting with regular expressions—they’re perfect for matching patterns like "starts with X, ends with Y, and can have messy characters in between". Let’s adjust your code to meet this requirement step by step.
Key Changes Explained
- Regex Pattern Matching: We’ll define regex patterns for each keyword to capture both exact spellings and "messy" variants (with special characters/typos).
- Count Accumulation: Your original code overwrote counts per file—we’ll switch to accumulating totals across all files instead.
- Dynamic Directory Usage: Fixed the hardcoded
/Settlementpath to respect the directory you pass via command line.
Modified Full Code
import os import glob import sys import re out1 = open("seen.txt", "w") out2 = open("missing.txt", "w") def count_words_in_dir(dirpath, word_patterns, action=None): # Initialize counts to 0 before processing files counts = {key: 0 for key in word_patterns.keys()} for filepath in glob.iglob(os.path.join(dirpath, '*.txt')): with open(filepath, 'r', encoding='utf-8') as f: data = f.read() for key, pattern in word_patterns.items(): # Find all matches for the regex pattern (case-insensitive) matches = re.findall(pattern, data, flags=re.IGNORECASE) counts[key] += len(matches) if action: action(filepath, counts.copy()) # Pass a copy to avoid overwriting during summary def print_summary(filepath, counts): for key, val in sorted(counts.items()): whichout = out1 if val > 0 else out2 print(filepath, file=whichout) print('{0}: {1}'.format(key, val), file=whichout) if __name__ == "__main__": if len(sys.argv) != 2: print("Usage: python script.py <directory_path>") sys.exit(1) target_dir = sys.argv[1] # Map keywords to their approximate match regex patterns word_patterns = { "annuity": r'\bann.*?y\b', "settlement": r'\bsett.*?nt\b' } count_words_in_dir(target_dir, word_patterns, action=print_summary) out1.close() out2.close()
Regex Pattern Breakdown
Take r'\bsett.*?nt\b' as an example:
\b: Matches a word boundary (ensures we don’t match partial words like "settlementxyz" or "abcsettlement")sett: The required prefix for "settlement" variants.*?: Matches any character (letters, symbols, numbers) in a non-greedy way—prevents matching across multiple wordsnt: The required suffixre.IGNORECASE: Makes matches case-insensitive, so "Settl#Ment" or "ANN&&ITY" will also be counted
Testing with Your Example Text
If your text file contains:
settlement settl#7*nt se##tl#ment ann&&ity annuity.
The script will count:
settlement: 3 (all three variants match the pattern)annuity: 2 (both variants match the pattern)
Customization Tips
- To restrict allowed middle characters (e.g., only letters and
#,*,&), adjust the regex tor'\bsett[a-zA-Z#*&]*nt\b'instead of.*?. - If you need to exclude certain characters, modify the regex character set to exclude them (e.g.,
[^0-9]to skip numbers).
内容的提问来源于stack exchange,提问作者JohnD
相关产品推荐
相关产品推荐

