You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将多个TXT文件内容分别存入DataFrame单行

单个TXT文件对应DataFrame一行的实现方案

你的问题出在使用pd.read_csv(sep='delimiter')时,会将文本文件的每一行拆分为DataFrame的单独一行,不符合「每个文件对应一行」的需求。要解决这个问题,直接读取整个文件的完整内容作为单个字符串即可,以下是优化后的代码:

import pandas as pd
import os

txtfolder = r'/home/spx072/Black_coffer_assignment/' # 修改为你的文件夹路径

# 批量收集文件名与对应完整内容
file_records = []
for root, _, files in os.walk(txtfolder):
    for file in files:
        if file.endswith('.txt'):
            full_path = os.path.join(root, file)
            # 读取整个文件内容为单个字符串
            with open(full_path, 'r', encoding='utf-8') as txt_file:
                full_content = txt_file.read()
            file_records.append({
                'Samplename': os.path.basename(file),
                'data': full_content
            })

# 一次性生成目标DataFrame
df = pd.DataFrame(file_records)
# 调整列顺序(按需选择)
df = df[['Samplename', 'data']]

关键说明

  1. 读取方式调整:用open().read()替代pd.read_csv,直接获取整个文件的完整内容,确保每个文件对应一个字符串值,不会拆分每行。
  2. 性能优化:先将所有文件数据存入列表,再一次性创建DataFrame,比循环拼接concat的效率更高,尤其适合处理大量文件。
  3. 编码兼容:添加encoding='utf-8'避免编码报错,若你的TXT文件使用其他编码(如GBK),可替换为对应编码值。

原代码问题解析

你之前用pd.read_csv(sep='delimiter', header=None)时,pandas会将文件的每一行识别为一条记录,因此每个文件的多行内容会被拆分成DataFrame的多行,这和「每个文件对应一行」的需求完全相反。

内容的提问来源于stack exchange,提问作者Abhishek Shahi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 23:01:03