You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python读取多TXT文件并将特定格式TXT转为DataFrame

Python读取TXT文件到DataFrame的两种场景解决方案

1. 批量读取多个TXT文件到单个DataFrame

如果目标TXT文件具有相同的列结构,可通过pandas+glob快速实现批量读取与合并:

import pandas as pd
import glob

# 匹配目标路径下所有TXT文件(示例为当前目录,可修改路径如"./data/*.txt")
txt_file_paths = glob.glob("*.txt")

# 循环读取并收集DataFrame
df_collection = []
for file_path in txt_file_paths:
    # 根据文件实际分隔符修改sep参数(如空格" "、制表符"\t"、逗号",")
    temp_df = pd.read_csv(file_path, sep="\t")
    # 可选:添加列标记数据来源文件
    temp_df["source_file"] = file_path.split("/")[-1]
    df_collection.append(temp_df)

# 合并所有DataFrame
final_df = pd.concat(df_collection, ignore_index=True)

若文件结构不一致,可在循环内针对不同文件单独处理列映射或筛选逻辑。


2. 读取键值对格式TXT为列名-行值的DataFrame

针对每行以[列名]: [值]格式存储的TXT文件,可解析为字典后转换为DataFrame:

单组数据场景(文件内仅一组键值对)

import pandas as pd

# 读取文件并过滤空行
with open("your_file.txt", "r", encoding="utf-8") as f:
    valid_lines = [line.strip() for line in f if line.strip()]

# 解析为键值对字典
data_dict = {}
for line in valid_lines:
    # 按第一个冒号拆分,避免值中含冒号导致错误
    col_name, col_value = line.split(":", 1)
    data_dict[col_name.strip()] = col_value.strip()

# 转换为DataFrame(一行数据对应所有列)
result_df = pd.DataFrame([data_dict])

多组数据场景(文件内有多组键值对,以空行分隔)

import pandas as pd

with open("your_file.txt", "r", encoding="utf-8") as f:
    # 按空行拆分数据组
    data_groups = f.read().split("\n\n")

data_list = []
for group in data_groups:
    lines = [line.strip() for line in group.split("\n") if line.strip()]
    group_dict = {}
    for line in lines:
        col_name, col_value = line.split(":", 1)
        group_dict[col_name.strip()] = col_value.strip()
    data_list.append(group_dict)

result_df = pd.DataFrame(data_list)

内容的提问来源于stack exchange,提问作者JoyanBhathena

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.21 07:27:23