如何用Python从多文本文件的大量字典中提取含指定K-V对的完整字典?
批量提取含特定键值对的Python字典方案
当然有超简便的实现方法!用Python写个小脚本就能高效搞定这300多个文件的批量处理,我给你详细拆解步骤和代码:
核心思路
其实逻辑很直白,就是三步走:
- 遍历目标文件夹里的所有文本文件
- 逐个解析文件里的Python字典,检查是否包含你要的键值对
- 把符合条件的完整字典写入结果文件
完整可运行代码示例
import os from pathlib import Path import ast # 用这个比eval更安全,推荐! # 先配置好你的参数,替换成实际值就行 TARGET_FOLDER = "./your_text_files_folder" # 你的目标文件夹路径 TARGET_KV = {"your_key": "your_target_value"} # 你要找的键值对 OUTPUT_FILE = "./matched_dicts_result.txt" # 结果保存的文件路径 def load_dicts_from_file(file_path): """从单个文本文件读取并解析所有Python字典""" dicts_list = [] with open(file_path, "r", encoding="utf-8") as f: for line_num, line in enumerate(f, 1): line = line.strip() if not line: continue try: # 用ast.literal_eval替代eval,避免执行恶意代码,更安全 dic = ast.literal_eval(line) if isinstance(dic, dict): dicts_list.append(dic) except (SyntaxError, ValueError): print(f"⚠️ 文件 {file_path} 第{line_num}行解析失败,跳过该行: {line}") continue return dicts_list def is_target_dict(dic, target_kv): """检查字典是否包含指定的所有键值对""" # 确保字典里的对应键值完全匹配 for key, expected_value in target_kv.items(): if dic.get(key) != expected_value: return False return True def main(): # 自动创建结果文件的父目录,避免路径不存在报错 Path(OUTPUT_FILE).parent.mkdir(parents=True, exist_ok=True) # 打开结果文件,准备写入 with open(OUTPUT_FILE, "w", encoding="utf-8") as output_f: # 遍历文件夹里的所有.txt文件(可根据实际后缀调整) for file_name in os.listdir(TARGET_FOLDER): file_path = os.path.join(TARGET_FOLDER, file_name) # 跳过非文件和非目标后缀的内容 if not os.path.isfile(file_path) or not file_name.lower().endswith(".txt"): continue print(f"🔄 正在处理文件: {file_name}") all_dicts = load_dicts_from_file(file_path) # 筛选符合条件的字典并写入 for dic in all_dicts: if is_target_dict(dic, TARGET_KV): # 把字典转成字符串,每行写一个,方便后续查看或再处理 output_f.write(f"{dic}\n") print(f"✅ 处理完成!所有符合条件的字典已保存到 {OUTPUT_FILE}") if __name__ == "__main__": main()
关键细节说明
- 安全性优先:我用了
ast.literal_eval()替代eval(),前者只会解析Python字面量(比如字典、列表这些),不会执行任意代码,哪怕文件里有奇怪内容也不会出问题,强烈推荐用这个! - 格式适配:默认假设每个字典单独占一行,如果你的文件里字典是用逗号分隔或者其他格式,只需要修改
load_dicts_from_file函数的解析逻辑就行——比如读取整个文件内容后,按分隔符拆分再逐个解析。 - 性能优化:如果觉得单线程处理300多个文件太慢,可以用Python的
concurrent.futures.ThreadPoolExecutor做并行处理,把文件处理任务分给多个线程,速度能提升不少。比如:from concurrent.futures import ThreadPoolExecutor # 修改main函数里的遍历部分为并行处理 def main(): Path(OUTPUT_FILE).parent.mkdir(parents=True, exist_ok=True) # 定义每个文件的处理函数 def process_file(file_path): print(f"🔄 正在处理文件: {os.path.basename(file_path)}") all_dicts = load_dicts_from_file(file_path) matched = [dic for dic in all_dicts if is_target_dict(dic, TARGET_KV)] return matched # 收集所有要处理的文件路径 file_paths = [] for file_name in os.listdir(TARGET_FOLDER): fp = os.path.join(TARGET_FOLDER, file_name) if os.path.isfile(fp) and file_name.lower().endswith(".txt"): file_paths.append(fp) # 用4个线程并行处理(可根据CPU核心数调整) with ThreadPoolExecutor(max_workers=4) as executor: results = executor.map(process_file, file_paths) # 把所有结果写入文件 with open(OUTPUT_FILE, "w", encoding="utf-8") as output_f: for matched_dicts in results: for dic in matched_dicts: output_f.write(f"{dic}\n") print(f"✅ 并行处理完成!结果已保存到 {OUTPUT_FILE}") - 容错性:代码里加了错误捕获,遇到解析失败的行会自动跳过并打印提示,不会因为个别坏行导致整个脚本崩溃。
内容的提问来源于stack exchange,提问作者funnyguy
相关产品推荐
相关产品推荐

