Jupyter Notebook调试Pandas大型DataFrame:定位格式错误行
定位CSV中格式错误的行(解决
ast.literal_eval报错) 问题背景
我有一个包含58k+行的CSV文件,其中record列存储的是字符串格式的可变长度字典。将文件读取为DataFrame后,执行以下代码尝试解析字典时触发错误:ValueError: malformed node or string on line 1: <ast.Name object at 0x17f1e73a0>,但小样本数据运行完全正常。需要定位出格式异常的具体行。
原处理代码:
df["record"] = df["record"].astype('str').apply(lambda x: x.replace("null", "None")) df['record']=df['record'].astype('str').apply(lambda x: ast.literal_eval(x)) df.record.apply(pd.Series)
可正常运行的样本DataFrame:
df=pd.DataFrame({'id': {0: 1689917349150, 1: 1689917349356, 2: 1689917349363, 3: 1689917349552, 4: 1689917349557, 5: 1689917349764, 6: 1689917349768}, 'record': {0: '{"tms_ProjectReleaseRelationship":{"project":{"uid":"myProj","source":"Jira"},"release":{"uid":"379000","source":"Jira"}}}', 1: '{"tms_Release":{"uid":"426949","name":"2016.","description":null,"startedAt":null,"releasedAt":null,"source":"Jira"}}', 2: '{"tms_ProjectReleaseRelationship":{"project":{"uid":"myProj","source":"Jira"},"release":{"uid":"123455","source":"Jira"}}}', 3: '{"tms_Release":{"uid":"426951","name":"2016.13","description":null,"startedAt":null,"releasedAt":null,"source":"Jira"}}', 4: '{"tms_ProjectReleaseRelationship":{"project":{"uid":"myProj","source":"Jira"},"release":{"uid":"426991","source":"Jira"}}}', 5: '{"tms_Release":{"uid":"427240","name":"2016","description":null,"startedAt":null,"releasedAt":null,"source":"Jira"}}', 6: '{"tms_ProjectReleaseRelationship":{"project":{"uid":"myProj","source":"Jira"},"release":{"uid":"427241","source":"Jira"}}}'}, 'index': {0: 56840, 1: 56841, 2: 56842, 3: 56843, 4: 56844, 5: 56845, 6: 56846})
定位错误行的解决方案
不要直接批量执行ast.literal_eval,改成逐行尝试并捕获异常,记录出错的行信息:
方法1:用apply封装安全解析函数
import ast def safe_parse(record_str): processed_str = str(record_str).replace("null", "None") try: return ast.literal_eval(processed_str) except Exception as e: # 返回错误标识和信息,方便后续筛选 return f"PARSE_ERROR: {str(e)}" # 对record列应用安全解析 df["parsed_record"] = df["record"].apply(safe_parse) # 筛选出解析失败的行 error_rows = df[df["parsed_record"].str.startswith("PARSE_ERROR:")] # 查看错误详情 print(f"共找到{len(error_rows)}条错误行:") print(error_rows[["id", "record", "parsed_record"]])
方法2:循环遍历逐行检查(更直观)
import ast error_info = [] for idx, row in df.iterrows(): record_str = str(row["record"]).replace("null", "None") try: ast.literal_eval(record_str) except Exception as e: error_info.append({ "index": idx, "id": row["id"], "record": row["record"], "error_msg": str(e) }) # 转换为DataFrame查看错误行 error_df = pd.DataFrame(error_info) print(error_df)
后续排查提示
定位到错误行后,重点检查record字段的格式问题:
- 是否存在未转义的双引号(比如字符串内部的
"没有用\"转义) - 是否有多余的逗号(比如字典末尾的逗号)
null替换是否彻底(比如存在大小写不一致的Null/NULL)- 是否有非标准的字典语法(比如单引号混用、缺少括号等)
内容的提问来源于stack exchange,提问作者Fariha Baloch
相关产品推荐
相关产品推荐

