Python中手动拼接的DataFrame文本无法解析为JSON的问题
问题分析与解决方案
错误根源
你的JSONDecodeError是因为手动拼接的JSON存在两个语法问题:
- 日期值
5/31/2022属于字符串,但未用双引号包裹,JSON语法强制要求所有字符串必须用双引号括起来; - Python中的
nan不是合法的JSON值,JSON中对应的空值是null,直接写入nan会导致解析失败。
修复方案
方案1:修正手动拼接的代码
修改你的函数,给日期添加双引号,并将nan替换为null:
import datetime import pandas as pd def combined_df_json(df): text = "{" for i in range(len(df)): # 处理Balance列的nan值 balance_val = str(df.Balance[i]) if pd.notna(df.Balance[i]) else "null" # 给日期字符串添加双引号 date_val = f'"{df.Date[i]}"' text += ( f'"{df.Tr_Id[i]}":' '{' f'"id": {df.Tr_Id[i]},' f'"date": {date_val},' f'"out": {df.Out[i]},' f'"in": {df.In[i]},' f'"balance": {"null" if balance_val == "nan" else f'"{balance_val}"'}' '},' ) # 移除最后一个多余的逗号并闭合大括号 text = text[:-1] + "}" return text
方案2:使用JSON工具生成(推荐)
手动拼接字符串极易出现语法错误,更可靠的方式是先构建Python字典,再用json模块生成合法JSON:
import json import pandas as pd def combined_df_json(df): result_dict = {} # 遍历DataFrame的每一行 for _, row in df.iterrows(): tr_id = str(row["Tr_Id"]) result_dict[tr_id] = { "id": row["Tr_Id"], "date": row["Date"], "out": row["Out"], "in": row["In"], "balance": row["Balance"] if pd.notna(row["Balance"]) else None } # json.dumps会自动处理字符串引号、null转换等问题 return json.dumps(result_dict)
方案3:直接用Pandas内置方法(最简)
Pandas自带to_json方法,一行代码就能完成转换,自动处理所有JSON语法规则:
# 将Tr_Id设为索引,以index格式输出JSON text_combined_csv = combined_df.set_index("Tr_Id").to_json(orient="index")
验证修复
使用上述任一方案生成的JSON文本,再执行你的加载代码:
import json with open(path+'cmbd_csv.json', "w") as f: f.write(text_combined_csv) with open(path+'cmbd_csv.json', "r") as f: file_contents = json.load(f)
此时不会再触发JSONDecodeError,可以正常解析JSON内容。
内容的提问来源于stack exchange,提问作者OpK
相关产品推荐
相关产品推荐

