如何在R中从百万行文本文件提取数据生成三列DataFrame
实现方案
处理思路
原始每行数据包含3个部分:开头的行号标记、中间带转义的JSON格式三元组、末尾的逗号和引号。我们只需要提取中间的JSON部分解析,再批量转换为DataFrame即可。该方案采用逐行读取的方式处理,内存占用低,适配百万级行的大文件场景。
代码实现
首先确保你已安装pandas依赖:pip install pandas
完整处理代码如下:
import pandas as pd import json # 替换为你的实际文件路径 file_path = "your_data.txt" processed_rows = [] with open(file_path, "r", encoding="utf-8") as f: for line in f: line = line.strip() if not line: continue # 提取有效JSON片段:拆分掉开头的行号标记、末尾的冗余符号 json_str = line.split('"', 1)[1].rsplit('",', 1)[0] # 解析JSON得到三个字段 p1, p2, year = json.loads(json_str) processed_rows.append((p1, p2, year)) # 生成目标DataFrame df = pd.DataFrame(processed_rows, columns=["Person1", "Person2", "Year"]) # 输出验证 print(df.head())
异常容错优化
如果你的文件存在格式异常的坏行,可以添加异常捕获逻辑跳过坏行,避免程序中断:
import pandas as pd import json file_path = "your_data.txt" processed_rows = [] bad_lines = [] with open(file_path, "r", encoding="utf-8") as f: for line in f: line = line.strip() if not line: continue try: json_str = line.split('"', 1)[1].rsplit('",', 1)[0] p1, p2, year = json.loads(json_str) processed_rows.append((p1, p2, year)) except Exception as e: bad_lines.append(line) df = pd.DataFrame(processed_rows, columns=["Person1", "Person2", "Year"]) print(f"处理完成,共生成{len(df)}条有效数据,跳过{len(bad_lines)}条异常行")
超大规模文件优化
如果文件大小超过内存容量,可以采用分块处理的方式,每处理N行就写入一次磁盘文件,避免内存溢出:
import pandas as pd import json file_path = "your_data.txt" chunk_size = 100000 # 每10万行存一次 chunk_idx = 0 current_chunk = [] with open(file_path, "r", encoding="utf-8") as f: for line in f: line = line.strip() if not line: continue try: json_str = line.split('"', 1)[1].rsplit('",', 1)[0] p1, p2, year = json.loads(json_str) current_chunk.append((p1, p2, year)) except: continue # 达到分块大小就写入磁盘 if len(current_chunk) >= chunk_size: df = pd.DataFrame(current_chunk, columns=["Person1", "Person2", "Year"]) df.to_csv(f"output_part_{chunk_idx}.csv", index=False) chunk_idx +=1 current_chunk = [] # 写入最后不足分块大小的剩余数据 if current_chunk: df = pd.DataFrame(current_chunk, columns=["Person1", "Person2", "Year"]) df.to_csv(f"output_part_{chunk_idx}.csv", index=False)
内容的提问来源于stack exchange,提问作者abdul samad
相关产品推荐
相关产品推荐

