You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Jupyter Notebook调试Pandas大型DataFrame:定位格式错误行

定位CSV中格式错误的行(解决ast.literal_eval报错)

问题背景

我有一个包含58k+行的CSV文件,其中record列存储的是字符串格式的可变长度字典。将文件读取为DataFrame后,执行以下代码尝试解析字典时触发错误:ValueError: malformed node or string on line 1: <ast.Name object at 0x17f1e73a0>,但小样本数据运行完全正常。需要定位出格式异常的具体行。

原处理代码:

df["record"] = df["record"].astype('str').apply(lambda x: x.replace("null", "None"))

df['record']=df['record'].astype('str').apply(lambda x: ast.literal_eval(x))
df.record.apply(pd.Series)

可正常运行的样本DataFrame:

df=pd.DataFrame({'id': {0: 1689917349150,
  1: 1689917349356,
  2: 1689917349363,
  3: 1689917349552,
  4: 1689917349557,
  5: 1689917349764,
  6: 1689917349768},
 'record': {0: '{"tms_ProjectReleaseRelationship":{"project":{"uid":"myProj","source":"Jira"},"release":{"uid":"379000","source":"Jira"}}}',
  1: '{"tms_Release":{"uid":"426949","name":"2016.","description":null,"startedAt":null,"releasedAt":null,"source":"Jira"}}',
  2: '{"tms_ProjectReleaseRelationship":{"project":{"uid":"myProj","source":"Jira"},"release":{"uid":"123455","source":"Jira"}}}',
  3: '{"tms_Release":{"uid":"426951","name":"2016.13","description":null,"startedAt":null,"releasedAt":null,"source":"Jira"}}',
  4: '{"tms_ProjectReleaseRelationship":{"project":{"uid":"myProj","source":"Jira"},"release":{"uid":"426991","source":"Jira"}}}',
  5: '{"tms_Release":{"uid":"427240","name":"2016","description":null,"startedAt":null,"releasedAt":null,"source":"Jira"}}',
  6: '{"tms_ProjectReleaseRelationship":{"project":{"uid":"myProj","source":"Jira"},"release":{"uid":"427241","source":"Jira"}}}'},
 'index': {0: 56840,
  1: 56841,
  2: 56842,
  3: 56843,
  4: 56844,
  5: 56845,
  6: 56846})

定位错误行的解决方案

不要直接批量执行ast.literal_eval,改成逐行尝试并捕获异常,记录出错的行信息:

方法1:用apply封装安全解析函数

import ast

def safe_parse(record_str):
    processed_str = str(record_str).replace("null", "None")
    try:
        return ast.literal_eval(processed_str)
    except Exception as e:
        # 返回错误标识和信息,方便后续筛选
        return f"PARSE_ERROR: {str(e)}"

# 对record列应用安全解析
df["parsed_record"] = df["record"].apply(safe_parse)

# 筛选出解析失败的行
error_rows = df[df["parsed_record"].str.startswith("PARSE_ERROR:")]

# 查看错误详情
print(f"共找到{len(error_rows)}条错误行:")
print(error_rows[["id", "record", "parsed_record"]])

方法2:循环遍历逐行检查(更直观)

import ast

error_info = []

for idx, row in df.iterrows():
    record_str = str(row["record"]).replace("null", "None")
    try:
        ast.literal_eval(record_str)
    except Exception as e:
        error_info.append({
            "index": idx,
            "id": row["id"],
            "record": row["record"],
            "error_msg": str(e)
        })

# 转换为DataFrame查看错误行
error_df = pd.DataFrame(error_info)
print(error_df)

后续排查提示

定位到错误行后,重点检查record字段的格式问题:

  • 是否存在未转义的双引号(比如字符串内部的"没有用\"转义)
  • 是否有多余的逗号(比如字典末尾的逗号)
  • null替换是否彻底(比如存在大小写不一致的Null/NULL)
  • 是否有非标准的字典语法(比如单引号混用、缺少括号等)

内容的提问来源于stack exchange,提问作者Fariha Baloch

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.15 18:02:44