You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python:在TXT文件中查找并统计单词的精确匹配与近似匹配

Solution for Approximate Keyword Matching in Text Files

Absolutely, this is totally achievable! The core fix here is swapping out exact string counting with regular expressions—they’re perfect for matching patterns like "starts with X, ends with Y, and can have messy characters in between". Let’s adjust your code to meet this requirement step by step.

Key Changes Explained

  • Regex Pattern Matching: We’ll define regex patterns for each keyword to capture both exact spellings and "messy" variants (with special characters/typos).
  • Count Accumulation: Your original code overwrote counts per file—we’ll switch to accumulating totals across all files instead.
  • Dynamic Directory Usage: Fixed the hardcoded /Settlement path to respect the directory you pass via command line.

Modified Full Code

import os
import glob
import sys
import re

out1 = open("seen.txt", "w")
out2 = open("missing.txt", "w")

def count_words_in_dir(dirpath, word_patterns, action=None):
    # Initialize counts to 0 before processing files
    counts = {key: 0 for key in word_patterns.keys()}
    for filepath in glob.iglob(os.path.join(dirpath, '*.txt')):
        with open(filepath, 'r', encoding='utf-8') as f:
            data = f.read()
            for key, pattern in word_patterns.items():
                # Find all matches for the regex pattern (case-insensitive)
                matches = re.findall(pattern, data, flags=re.IGNORECASE)
                counts[key] += len(matches)
        if action:
            action(filepath, counts.copy())  # Pass a copy to avoid overwriting during summary

def print_summary(filepath, counts):
    for key, val in sorted(counts.items()):
        whichout = out1 if val > 0 else out2
        print(filepath, file=whichout)
        print('{0}: {1}'.format(key, val), file=whichout)

if __name__ == "__main__":
    if len(sys.argv) != 2:
        print("Usage: python script.py <directory_path>")
        sys.exit(1)
    target_dir = sys.argv[1]
    # Map keywords to their approximate match regex patterns
    word_patterns = {
        "annuity": r'\bann.*?y\b',
        "settlement": r'\bsett.*?nt\b'
    }
    count_words_in_dir(target_dir, word_patterns, action=print_summary)
    out1.close()
    out2.close()

Regex Pattern Breakdown

Take r'\bsett.*?nt\b' as an example:

  • \b: Matches a word boundary (ensures we don’t match partial words like "settlementxyz" or "abcsettlement")
  • sett: The required prefix for "settlement" variants
  • .*?: Matches any character (letters, symbols, numbers) in a non-greedy way—prevents matching across multiple words
  • nt: The required suffix
  • re.IGNORECASE: Makes matches case-insensitive, so "Settl#Ment" or "ANN&&ITY" will also be counted

Testing with Your Example Text

If your text file contains:

settlement settl#7*nt se##tl#ment ann&&ity annuity.

The script will count:

  • settlement: 3 (all three variants match the pattern)
  • annuity: 2 (both variants match the pattern)

Customization Tips

  • To restrict allowed middle characters (e.g., only letters and #, *, &), adjust the regex to r'\bsett[a-zA-Z#*&]*nt\b' instead of .*?.
  • If you need to exclude certain characters, modify the regex character set to exclude them (e.g., [^0-9] to skip numbers).

内容的提问来源于stack exchange,提问作者JohnD

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 13:19:08