You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python遍历目录检索关键词 CSV导出仅存最后一条如何追加写入

CSV文件写入仅保留最后一条记录的修复方案

问题根因

  • 代码中赋值给df的是拼接后的普通字符串,并非pandas DataFrame对象,直接调用to_csv()本身就会触发属性错误,无法正常执行。
  • to_csv()方法默认使用覆盖写入模式(mode='w'),每次调用都会清空目标文件的原有内容再写入当前数据,循环中反复调用会不断覆盖上一次的写入结果,最终仅保留最后一次循环生成的记录。
  • Windows路径直接写E:\KeySearch.txt存在转义风险,反斜杠\在Python字符串中是转义字符,可能导致路径识别失败。
  • 直接拼接逗号生成CSV内容的写法存在格式风险:如果匹配到的行内容本身包含逗号,会出现CSV列错位的问题。

修复代码(推荐:批量收集后一次性写入)

这种方式IO开销最低,性能最好,适合绝大多数场景:

import os
import pandas as pd

words = ['Bonus Allocation', 'Benefits', 'bird']
# 初始化列表存储所有匹配结果
result = []

for root, _, files in os.walk(r'E:\CSV_DOCS'):
    for path in filter(lambda p: p.endswith('.csv'), files):
        file_full_path = os.path.join(root, path)
        # 读取文件时指定编码,避免乱码
        with open(file_full_path, encoding='utf-8') as f:
            # 行号从1开始计数,和原有输出逻辑对齐
            for i, line in enumerate(f.readlines(), start=1):
                line_stripped = line.strip()
                for word in filter(lambda w: w in line_stripped, words):
                    print(f'{path}, {i}, {word}, {line_stripped}')
                    # 按结构化字段存储结果,避免逗号错位
                    result.append({
                        '文件名': path,
                        '行号': i,
                        '匹配关键词': word,
                        '行内容': line_stripped
                    })

# 所有遍历结束后一次性写入文件
df = pd.DataFrame(result)
# 路径用原始字符串避免转义问题,index=False关闭多余的行索引列
# utf-8-sig编码支持Excel直接打开不乱码,需要存txt直接改后缀即可
df.to_csv(r'E:\KeySearch.csv', index=False, encoding='utf-8-sig')

可选方案:循环内追加写入

如果匹配结果量级极大,不想把所有结果暂存在内存中,可以在循环内使用追加模式写入,注意只在第一次写入时输出表头,后续追加跳过表头:

import os
import pandas as pd

words = ['Bonus Allocation', 'Benefits', 'bird']
output_path = r'E:\KeySearch.csv'
# 初始化标记:判断文件是否已存在,不存在则第一次写入时加表头
file_exists = os.path.exists(output_path)

for root, _, files in os.walk(r'E:\CSV_DOCS'):
    for path in filter(lambda p: p.endswith('.csv'), files):
        file_full_path = os.path.join(root, path)
        with open(file_full_path, encoding='utf-8') as f:
            for i, line in enumerate(f.readlines(), start=1):
                line_stripped = line.strip()
                for word in filter(lambda w: w in line_stripped, words):
                    print(f'{path}, {i}, {word}, {line_stripped}')
                    temp_df = pd.DataFrame([{
                        '文件名': path,
                        '行号': i,
                        '匹配关键词': word,
                        '行内容': line_stripped
                    }])
                    # 指定追加模式a,控制表头仅写入一次
                    temp_df.to_csv(
                        output_path,
                        mode='a',
                        header=not file_exists,
                        index=False,
                        encoding='utf-8-sig'
                    )
                    file_exists = True

内容的提问来源于stack exchange,提问作者William Smith

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.29 20:27:16