You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas行转列提取数据时pivot提示索引含重复条目无法重塑如何解决

解决方案

报错原因

你调用pivot时没有指定行索引参数,pandas默认使用原DataFrame的索引作为行标识,而你执行explode生成的3行数据索引都是初始值0,存在重复,所以触发了「Index contains duplicate entries, cannot reshape」的报错。

单指标场景(你当前的使用场景)

你当前所有行的name都是固定的page_fans,不需要用透视操作,直接拆分values列的字典即可得到目标结构,完整可运行代码如下:

import pandas as pd

row =  {'name': 'page_fans', 'period': 'day', 'values': [{'value': 111, 'end_time': '2021-09-13T07:00:00+0000'}, {'value': 233, 'end_time': '2021-09-14T07:00:00+0000'}, {'value': 551, 'end_time': '2021-09-15T07:00:00+0000'}], 'title': 'Lifetime Total Likes', 'description': 'Lifetime: The total number of people who have liked your Page. (Unique Users)', 'id': '247111/insights/page_fans/day'}

pat_id = r'(\d+)'
df = pd.json_normalize(row)
df['id'] = df['id'].astype(str).str.extract(pat_id)
df = df.explode('values').reset_index(drop=True)

# 拆分values列的字典为单独字段
df[['page_fans', 'end_time']] = df['values'].apply(pd.Series)
# 筛选保留需要的列即可得到目标结构
result = df[['page_fans', 'end_time', 'id']]

多指标通用场景

如果你的实际数据中name列存在多个不同的指标名,需要更通用的转换方案,可以用如下代码:

# 先拆分values字段
df[['value', 'end_time']] = df['values'].apply(pd.Series)
# 新增辅助行索引避免pivot时索引重复
df['row_num'] = df.groupby(['name', 'id']).cumcount()
# 透视转换
pivot_df = df.pivot(index=['row_num', 'id', 'end_time'], columns='name', values='value').reset_index()
# 清理多余列得到最终结果
result = pivot_df.drop(columns='row_num').rename_axis(columns=None)

内容的提问来源于stack exchange,提问作者KristiLuna

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.04 03:48:00