如何用Python统计文件中表达式出现次数及修复JSON统计错误
解决正则匹配次数统计错误的问题
看起来你已经搞定了文件读取和匹配内容高亮的部分,很棒!现在问题出在统计匹配次数时误把文件行数当成了实际匹配数——这个坑我以前也踩过,大概率是你在处理文件时,不小心把遍历行的计数器当成了匹配次数。
问题根源
你可能是在逐行读取文件时,用了类似enumerate的行号变量来统计“次数”,但行号只是文件的行数,和正则匹配的实际次数完全无关。正确的做法应该是统计re.findall返回的匹配列表的长度,或者逐行累加每一行的匹配数量。
修正方案
下面是完整的修正代码,包含高亮显示和正确统计匹配次数并保存JSON的逻辑:
import re import json def highlight_matches(text, pattern): # 用ANSI红色高亮匹配内容(终端中生效) return re.sub(pattern, lambda match: f"\033[91m{match.group()}\033[0m", text) def process_file(file_path, regex_pattern): try: # 一次性读取文件内容(小文件适用) with open(file_path, 'r', encoding='utf-8') as file: file_content = file.read() # 获取所有匹配结果,计算准确次数 all_matches = re.findall(regex_pattern, file_content) total_matches = len(all_matches) # 输出高亮内容 print("高亮后的文件内容:\n") print(highlight_matches(file_content, regex_pattern)) # 准备要存入JSON的结果数据 result_info = { "文件名": file_path, "匹配表达式": regex_pattern, "出现次数": total_matches } # 写入JSON文件,格式化输出更易读 with open("match_statistics.json", 'w', encoding='utf-8') as json_file: json.dump(result_info, json_file, ensure_ascii=False, indent=4) print(f"\n✅ 统计结果已保存到match_statistics.json,共匹配到{total_matches}次") except FileNotFoundError: print(f"❌ 错误:找不到文件 {file_path}") except re.error: print("❌ 错误:正则表达式格式无效,请检查后重试") if __name__ == "__main__": target_file = input("请输入目标文件名:") user_pattern = input("请输入要匹配的正则表达式:") process_file(target_file, user_pattern)
针对大文件的优化(逐行处理)
如果你的文件很大,不适合一次性读入内存,可以改成逐行处理并累加匹配次数:
# 替换process_file函数中的读取和统计部分 match_count = 0 highlighted_content = [] with open(file_path, 'r', encoding='utf-8') as file: for line in file: # 统计当前行的匹配数并累加 line_matches = re.findall(regex_pattern, line) match_count += len(line_matches) # 高亮当前行并加入列表 highlighted_content.append(highlight_matches(line, regex_pattern)) # 合并高亮内容 highlighted_content = ''.join(highlighted_content)
关键注意点
- 绝对不要用文件的行数来代替匹配次数,行数和匹配数没有直接关系(一行可能有多个匹配,也可能没有)
re.findall返回的是所有匹配字符串组成的列表,len()就是准确的总匹配数- 如果需要区分非重叠匹配/重叠匹配,可能需要用
re.finditer,但大部分场景下findall足够用
内容的提问来源于stack exchange,提问作者Py Dev
相关产品推荐
相关产品推荐

