You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas:从另一个DataFrame匹配最接近的分组数据

Pandas DataFrame 扁平化处理方案

假设你已经得到包含x_time、x_price(来自df_mapped)、original_price、price_diff(差值)以及df_original其他字段的中间DataFrame,以下是完成扁平化的具体步骤:

步骤1:分组筛选前3个最小差值行

按x_time和df_mapped的唯一标识(比如原df_mapped的索引,或自定义主键列)分组,对price_diff升序排序后取前3行:

# 假设中间DataFrame名为df_interim,且包含df_mapped的索引列(命名为mapped_idx)
df_top3 = df_interim.groupby(['x_time', 'mapped_idx']).apply(
    lambda x: x.sort_values('price_diff').head(3)
).reset_index(drop=True)

步骤2:为分组内的行添加序号

给每个分组内的前3行标记1-3的序号,为后续扁平化做准备:

df_top3['rank'] = df_top3.groupby(['x_time', 'mapped_idx']).cumcount() + 1

步骤3: pivot实现多行转单行

将分组后的多行数据扁平化为单行,把df_original的字段按序号展开:

# 替换成你需要展开的df_original字段,比如original_price、col_a、col_b
df_flat = df_top3.pivot(
    index=['x_time', 'mapped_idx'],
    columns='rank',
    values=['original_price', 'col_a', 'col_b']
)

# 重命名列名,让格式更直观(例如original_price_1、col_a_2)
df_flat.columns = ['_'.join(map(str, col)) for col in df_flat.columns]
df_flat = df_flat.reset_index()

步骤4:合并回原df_mapped

把扁平化后的结果与df_mapped关联,得到最终结果:

final_df = df_mapped.merge(df_flat, on=['x_time', 'mapped_idx'], how='left')

如果你的中间DataFrame未保留df_mapped的唯一标识,可在前期合并阶段就将df_mapped的索引或主键带入,确保分组时能精准对应到原df_mapped的每一行。


内容的提问来源于stack exchange,提问作者user20787745

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 10:45:42