Python遍历目录检索关键词 CSV导出仅存最后一条如何追加写入
CSV文件写入仅保留最后一条记录的修复方案
问题根因
- 代码中赋值给
df的是拼接后的普通字符串,并非pandas DataFrame对象,直接调用to_csv()本身就会触发属性错误,无法正常执行。 to_csv()方法默认使用覆盖写入模式(mode='w'),每次调用都会清空目标文件的原有内容再写入当前数据,循环中反复调用会不断覆盖上一次的写入结果,最终仅保留最后一次循环生成的记录。- Windows路径直接写
E:\KeySearch.txt存在转义风险,反斜杠\在Python字符串中是转义字符,可能导致路径识别失败。 - 直接拼接逗号生成CSV内容的写法存在格式风险:如果匹配到的行内容本身包含逗号,会出现CSV列错位的问题。
修复代码(推荐:批量收集后一次性写入)
这种方式IO开销最低,性能最好,适合绝大多数场景:
import os import pandas as pd words = ['Bonus Allocation', 'Benefits', 'bird'] # 初始化列表存储所有匹配结果 result = [] for root, _, files in os.walk(r'E:\CSV_DOCS'): for path in filter(lambda p: p.endswith('.csv'), files): file_full_path = os.path.join(root, path) # 读取文件时指定编码,避免乱码 with open(file_full_path, encoding='utf-8') as f: # 行号从1开始计数,和原有输出逻辑对齐 for i, line in enumerate(f.readlines(), start=1): line_stripped = line.strip() for word in filter(lambda w: w in line_stripped, words): print(f'{path}, {i}, {word}, {line_stripped}') # 按结构化字段存储结果,避免逗号错位 result.append({ '文件名': path, '行号': i, '匹配关键词': word, '行内容': line_stripped }) # 所有遍历结束后一次性写入文件 df = pd.DataFrame(result) # 路径用原始字符串避免转义问题,index=False关闭多余的行索引列 # utf-8-sig编码支持Excel直接打开不乱码,需要存txt直接改后缀即可 df.to_csv(r'E:\KeySearch.csv', index=False, encoding='utf-8-sig')
可选方案:循环内追加写入
如果匹配结果量级极大,不想把所有结果暂存在内存中,可以在循环内使用追加模式写入,注意只在第一次写入时输出表头,后续追加跳过表头:
import os import pandas as pd words = ['Bonus Allocation', 'Benefits', 'bird'] output_path = r'E:\KeySearch.csv' # 初始化标记:判断文件是否已存在,不存在则第一次写入时加表头 file_exists = os.path.exists(output_path) for root, _, files in os.walk(r'E:\CSV_DOCS'): for path in filter(lambda p: p.endswith('.csv'), files): file_full_path = os.path.join(root, path) with open(file_full_path, encoding='utf-8') as f: for i, line in enumerate(f.readlines(), start=1): line_stripped = line.strip() for word in filter(lambda w: w in line_stripped, words): print(f'{path}, {i}, {word}, {line_stripped}') temp_df = pd.DataFrame([{ '文件名': path, '行号': i, '匹配关键词': word, '行内容': line_stripped }]) # 指定追加模式a,控制表头仅写入一次 temp_df.to_csv( output_path, mode='a', header=not file_exists, index=False, encoding='utf-8-sig' ) file_exists = True
内容的提问来源于stack exchange,提问作者William Smith
相关产品推荐
相关产品推荐

