You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在Google Colab中把Pandas数据集每行文本转为独立TXT文件

解决方案

步骤说明与代码实现

1. 环境确认

Google Colab默认已预装pandas和pyarrow(读取feather格式文件的必要依赖),无需额外安装。若读取时报错,可执行以下命令补装:

!pip install pandas pyarrow

2. 完整代码

import pandas as pd
import os

# 读取zstd压缩的feather数据集
df = pd.read_feather("/content/preprocessed_sample.ftr.zstd")

# 创建目标文件夹,已存在则跳过
os.makedirs("examples", exist_ok=True)

# 遍历每行生成独立文本文件
for line_idx, row in df.iterrows():
    # 生成line1.txt、line2.txt格式的文件名
    file_path = f"examples/line{line_idx + 1}.txt"
    # 写入文本内容,指定utf-8编码避免乱码
    with open(file_path, "w", encoding="utf-8") as f:
        f.write(row["text"])

print("所有文件已成功保存至examples文件夹")

关键细节

  • 若DataFrame的index列是自定义的行序号(而非默认从0开始的索引),将line_idx + 1替换为row['index'],即可让文件名与自定义索引匹配
  • 写入时指定utf-8编码,可避免中文、特殊字符出现乱码
  • Colab中上传的文件默认存放在/content/目录下,若文件路径不同,需修改pd.read_feather中的路径参数

内容的提问来源于stack exchange,提问作者John Angelopoulos

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.02 07:40:42