如何将子数组itemdata与主数组saledata匹配关联?
主数组与子数组的匹配合并解决方案
现有数据说明
主数组 saledata 结构:
id sub-array 0 001 [{'type': 'line_items', 'id': '78', 'attributes': {'status': 'allocated', 'quantity': 1, 'various_other_data': 'etc'}}] 1 002 [{'type': 'line_items', 'id': '80', 'attributes': {'status': 'allocated', 'quantity': 2, 'various_other_data': 'etc'}}] 2 003 [{'type': 'line_items', 'id': '85', 'attributes': {'status': 'allocated', 'quantity': 1, 'various_other_data': 'etc'}}, {'type': 'line_items', 'id': '86', 'attributes': {'status': 'allocated', 'quantity': 1, 'various_other_data': 'etc'}}] 3 004 [{'type': 'line_items', 'id': '92', 'attributes': {'status': 'allocated', 'quantity': 2, 'various_other_data': 'etc'}}, {'type': 'line_items', 'id': '93', 'attributes': {'status': 'allocated', 'quantity': 2, 'various_other_data': 'etc'}}]
子数组 itemdata 结构(已归一化):
type id attributes.status attributes.quantity attributes.various_other_data 0 line_item 78 allocated 1 etc 0 line_item 80 allocated 2 etc 0 line_item 85 allocated 1 etc 1 line_item 86 allocated 1 etc 0 line_item 92 allocated 2 etc 1 line_item 93 allocated 2 etc
错误原因说明
你之前使用 if df['sub-array'].str.contains(f) == True 触发"Series真值判断模糊"错误,是因为 str.contains 返回的是布尔型Series,直接用if判断整个Series时,Pandas无法确定你要判断所有元素为True还是存在至少一个True,因此抛出警告。
可行解决方案
方案1:直接从主数组展开合并(推荐)
无需提前归一化itemdata,直接对saledata的sub-array列展开并解析,一步得到目标结果:
import pandas as pd # 展开saledata的sub-array列,保留原行索引 saledata_expanded = saledata.explode('sub-array', ignore_index=False) # 将展开后的字典数据归一化为DataFrame sub_array_df = pd.json_normalize(saledata_expanded['sub-array']) # 合并原saledata的id列与归一化后的子数组数据 result = pd.concat([saledata_expanded[['id']], sub_array_df], axis=1) # 调整列名以匹配期望格式 result = result.rename(columns={ 'type': 'type', 'id': 'itemdata.id', 'attributes.status': 'itemdata.attributes.status', 'attributes.quantity': 'itemdata.attributes.quantity', 'attributes.various_other_data': 'itemdata.attributes.various_other_data' }) # 可选:重置索引并调整格式 result = result.reset_index(drop=False).rename(columns={'index': ''})
方案2:基于已有的itemdata进行匹配
如果你已经有归一化后的itemdata,可以先建立主数组id与子项id的映射关系,再合并:
import pandas as pd # 展开saledata并提取主id与子项id的映射 saledata_expanded = saledata.explode('sub-array') saledata_expanded['item_id'] = saledata_expanded['sub-array'].apply(lambda x: x['id']) id_mapping = saledata_expanded[['id', 'item_id']] # 合并itemdata与映射表 result = pd.merge(itemdata, id_mapping, left_on='id', right_on='item_id') # 调整列顺序和名称到期望格式 result = result[['id', 'type', 'id_x', 'attributes.status', 'attributes.quantity']].rename(columns={ 'id_x': 'itemdata.id', 'attributes.status': 'itemdata.attributes.status', 'attributes.quantity': 'itemdata.attributes.quantity' })
最终结果
执行上述代码后,将得到符合需求的合并结果:
id type itemdata.id itemdata.attributes.status itemdata.attributes.quantity 0 001 line_items 78 allocated etc 1 002 line_items 80 allocated etc 2 003 line_items 85 allocated etc 2 003 line_items 86 allocated etc 3 004 line_items 92 allocated etc 3 004 line_items 93 allocated etc
内容的提问来源于stack exchange,提问作者notanothercliche
相关产品推荐
相关产品推荐

