Python关键词搜索工具无结果且耗时过长问题求助
问题排查与解决方案
一、核心问题排查
1. 关键词读取验证
先确认脚本是否正确读取keywords.txt,避免编码或空白字符问题:
with open('keywords.txt', 'r', encoding='utf-8') as f: # 过滤空行和首尾空白 keywords = [line.strip() for line in f if line.strip()] # 打印调试,确认关键词已读入 print(f"已加载关键词:{keywords}")
2. 遍历效率优化
耗时极长大概率是因为无差别遍历了所有文件/目录,跳过无关内容:
import os # 定义要排除的目录和非文本文件后缀 EXCLUDE_DIRS = {'.git', 'node_modules', '__pycache__', 'venv'} EXCLUDE_EXTS = {'.exe', '.png', '.jpg', '.zip', '.rar'} def traverse_target_dir(root_dir): all_files = [] for root, dirs, files in os.walk(root_dir): # 移除要排除的目录,避免递归进入 dirs[:] = [d for d in dirs if d not in EXCLUDE_DIRS] for file in files: file_ext = os.path.splitext(file)[1].lower() if file_ext not in EXCLUDE_EXTS: all_files.append(os.path.join(root, file)) return all_files
3. 匹配逻辑修正
检查是否存在匹配逻辑错误,比如大小写不敏感、匹配符误用:
def check_keyword_match(file_content, keywords): content_lower = file_content.lower() # 统一转为小写,避免大小写遗漏 return any(keyword.lower() in content_lower for keyword in keywords)
二、生成目录结构式输出
要输出类似目录树的结果,需按层级整理文件并标记匹配项:
def build_directory_tree(root_dir, matched_files): tree_lines = [] for root, dirs, files in os.walk(root_dir): # 计算当前目录层级,生成缩进 depth = root.replace(root_dir, '').count(os.sep) indent = '│ ' * depth + '├── ' # 添加目录节点 tree_lines.append(f"{indent}{os.path.basename(root)}/") file_indent = '│ ' * (depth + 1) + '├── ' for file in files: full_path = os.path.join(root, file) # 用[✓]标记匹配的文件,[ ]标记未匹配 status = '[✓] ' if full_path in matched_files else '[ ] ' tree_lines.append(f"{file_indent}{status}{file}") return '\n'.join(tree_lines) # 最后写入output.txt with open('output.txt', 'w', encoding='utf-8') as f: if matched_files: f.write(build_directory_tree(target_directory, matched_files)) else: f.write("未找到包含关键词的文件")
三、提速优化
如果文件数量极大,用线程池并行处理文件读取和匹配:
from concurrent.futures import ThreadPoolExecutor def process_single_file(file_path, keywords): try: # 忽略编码错误,避免因个别文件中断 with open(file_path, 'r', encoding='utf-8', errors='ignore') as f: content = f.read() if check_keyword_match(content, keywords): return file_path except Exception: # 跳过无法读取的文件 return None # 调用示例 target_files = traverse_target_dir(target_directory) matched_files = set() with ThreadPoolExecutor(max_workers=4) as executor: results = executor.map(lambda path: process_single_file(path, keywords), target_files) matched_files = {res for res in results if res is not None}
内容的提问来源于stack exchange,提问作者Дима Пузырин
相关产品推荐
相关产品推荐

