Python脚本无法匹配搜索词列表全部条目问题求助
多搜索词匹配不全问题的分析与修复
问题根源
你的脚本存在关键逻辑错误:遍历搜索词列表时,找到第一个匹配条目后就执行了break语句。这会导致:
- 单个文件即使包含多个搜索词,只会检测并记录列表中最先出现的那个匹配词,后续搜索词不会再对该文件进行检查。
- 如果某个搜索词(比如
ERR:)在列表中位置靠后,且所有包含它的文件同时也包含列表中排在它前面的搜索词,那么这个词永远不会被检测到,自然不会出现在记录里。 - 同时,同一个文件会因为匹配多个搜索词被重复复制,造成冗余。
修复方案
修改搜索逻辑,去掉break,同时避免重复复制文件和重复记录(可选):
import os import shutil import datetime source_folder = input("Enter the source folder path: ") search_text_list = ["x exception", "Displayed", "!!!!!", "thermal event", "ERR:", "WRN", "InstrumentMonitorEvent"] target_folder_name = "NovaSeq 6000 Parsing" match_file_folder_name = "NovaSeq 6000 Analyzer Output_" + str(datetime.datetime.now().strftime("%Y-%m-%d %H-%M-%S")) match_file_info = "Matched Search Terms.txt" desktop = os.path.join(os.path.join(os.environ['USERPROFILE']), 'Desktop') target_folder = os.path.join(desktop, target_folder_name) match_file_folder = os.path.join(target_folder, match_file_folder_name) match_file_path = os.path.join(match_file_folder, match_file_info) if not os.path.exists(target_folder): os.makedirs(target_folder) if not os.path.exists(match_file_folder): os.makedirs(match_file_folder) matched_search_terms = [] copied_files = set() # 记录已复制的文件路径,避免重复操作 for root, dirs, files in os.walk(source_folder): if "ETF" in dirs: dirs.remove("ETF") for file in files: # 简化跳过文件的判断逻辑 skip_keywords = ["Warnings_And_Errors", "RunSetup", "Wash"] if any(keyword in file for keyword in skip_keywords): continue full_path = os.path.join(root, file) with open(full_path, 'r', errors='ignore') as f: file_content = f.read().lower() # 收集当前文件匹配到的所有搜索词 file_matches = [term for term in search_text_list if term.lower() in file_content] if file_matches: # 将当前文件的所有匹配词加入总列表 matched_search_terms.extend(file_matches) # 仅复制未处理过的文件 if full_path not in copied_files: shutil.copy2(full_path, match_file_folder) copied_files.add(full_path) # 可选:对匹配词去重,确保每个词只记录一次(保持原搜索顺序) matched_search_terms = list(dict.fromkeys(matched_search_terms)) with open(match_file_path, 'w') as f: if matched_search_terms: f.write("\n".join(matched_search_terms)) else: f.write("NA")
关键修改点
- 移除
break语句:让每个文件完整遍历所有搜索词,找到全部匹配条目 - 新增
copied_files集合:避免同一个文件因匹配多个搜索词被重复复制 - 简化跳过文件的判断:用
any()替代多个if continue,代码更简洁 - 可选去重处理:用
dict.fromkeys()在保持原搜索词顺序的前提下去重,避免同一词多次记录
内容的提问来源于stack exchange,提问作者Brad Billstein
相关产品推荐
相关产品推荐

