You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R中从百万行文本文件提取数据生成三列DataFrame

实现方案

处理思路

原始每行数据包含3个部分:开头的行号标记、中间带转义的JSON格式三元组、末尾的逗号和引号。我们只需要提取中间的JSON部分解析,再批量转换为DataFrame即可。该方案采用逐行读取的方式处理,内存占用低,适配百万级行的大文件场景。

代码实现

首先确保你已安装pandas依赖:
pip install pandas
完整处理代码如下:

import pandas as pd
import json

# 替换为你的实际文件路径
file_path = "your_data.txt"
processed_rows = []

with open(file_path, "r", encoding="utf-8") as f:
    for line in f:
        line = line.strip()
        if not line:
            continue
        # 提取有效JSON片段:拆分掉开头的行号标记、末尾的冗余符号
        json_str = line.split('"', 1)[1].rsplit('",', 1)[0]
        # 解析JSON得到三个字段
        p1, p2, year = json.loads(json_str)
        processed_rows.append((p1, p2, year))

# 生成目标DataFrame
df = pd.DataFrame(processed_rows, columns=["Person1", "Person2", "Year"])

# 输出验证
print(df.head())

异常容错优化

如果你的文件存在格式异常的坏行,可以添加异常捕获逻辑跳过坏行,避免程序中断:

import pandas as pd
import json

file_path = "your_data.txt"
processed_rows = []
bad_lines = []

with open(file_path, "r", encoding="utf-8") as f:
    for line in f:
        line = line.strip()
        if not line:
            continue
        try:
            json_str = line.split('"', 1)[1].rsplit('",', 1)[0]
            p1, p2, year = json.loads(json_str)
            processed_rows.append((p1, p2, year))
        except Exception as e:
            bad_lines.append(line)

df = pd.DataFrame(processed_rows, columns=["Person1", "Person2", "Year"])
print(f"处理完成,共生成{len(df)}条有效数据,跳过{len(bad_lines)}条异常行")

超大规模文件优化

如果文件大小超过内存容量,可以采用分块处理的方式,每处理N行就写入一次磁盘文件,避免内存溢出:

import pandas as pd
import json

file_path = "your_data.txt"
chunk_size = 100000 # 每10万行存一次
chunk_idx = 0
current_chunk = []

with open(file_path, "r", encoding="utf-8") as f:
    for line in f:
        line = line.strip()
        if not line:
            continue
        try:
            json_str = line.split('"', 1)[1].rsplit('",', 1)[0]
            p1, p2, year = json.loads(json_str)
            current_chunk.append((p1, p2, year))
        except:
            continue
        # 达到分块大小就写入磁盘
        if len(current_chunk) >= chunk_size:
            df = pd.DataFrame(current_chunk, columns=["Person1", "Person2", "Year"])
            df.to_csv(f"output_part_{chunk_idx}.csv", index=False)
            chunk_idx +=1
            current_chunk = []
# 写入最后不足分块大小的剩余数据
if current_chunk:
    df = pd.DataFrame(current_chunk, columns=["Person1", "Person2", "Year"])
    df.to_csv(f"output_part_{chunk_idx}.csv", index=False)

内容的提问来源于stack exchange,提问作者abdul samad

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.25 13:15:04