Python如何将DataFrame列中的JSON列表拆分生成多行新数据
解决方案
核心处理逻辑:你的custbody_attachment_1列存储的是JSON数组格式字符串,同时存在空值、异常格式内容,按以下步骤处理即可:
完整可运行代码
import pandas as pd import json # 自定义JSON数组解析函数,兼容空值、异常格式场景 def parse_json_arr(s): # 处理空值、空白字符串场景 if pd.isna(s) or str(s).strip() == '': return [{}] try: # 可根据实际数据异常补充清洗规则,示例为删除字段开头多余的数字 cleaned_s = str(s).strip().lstrip('0123456789') return json.loads(cleaned_s) except: # 解析失败返回空字典,避免程序中断 return [{}] # 1. 将附件列转换为Python列表格式 lineitem['json_list'] = lineitem['custbody_attachment_1'].apply(parse_json_arr) # 2. 把列表拆分为多行 lineitem = lineitem.explode('json_list', ignore_index=True) # 3. 提取FL、NM字段为单独列 lineitem['FL'] = lineitem['json_list'].apply(lambda x: x.get('FL', '')) lineitem['NM'] = lineitem['json_list'].apply(lambda x: x.get('NM', '')) # 4. 删除中间列和原附件列,得到最终结果 res_df = lineitem.drop(columns=['custbody_attachment_1', 'json_list']) # 输出查看结果 print(res_df)
如果遇到其他JSON解析报错,只需要在parse_json_arr函数里补充对应的字符串清洗规则即可。
内容的提问来源于stack exchange,提问作者Ilya.s
相关产品推荐
相关产品推荐

