You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Pandas中按匹配字段高效合并个体与州级DataFrame?

Pandas高效合并个体与州级数据方案

步骤1:将taxframe转换为宽格式

taxframe是长格式存储(每个指标单独一行),先把它转成宽格式,让Total Tax/Pack和Avg Cost/Pack成为独立列,才能和个体级数据匹配:

# 先清理重复值,确保每个Location+Year+SubMeasure组合唯一
taxframe_clean = taxframe.drop_duplicates(subset=['Location', 'Year', 'SubMeasure'])

# 透视转换为宽格式
taxframe_wide = taxframe_clean.pivot(
    index=['Location', 'Year'],
    columns='SubMeasure',
    values='Value'
).reset_index()

# 可选:替换列名中的空格,避免后续操作报错
taxframe_wide.columns = [col.replace(' ', '_') for col in taxframe_wide.columns]

步骤2:按Location和Year合并数据

用Pandas内置的merge方法完成合并——这是向量化操作,底层由C实现,效率远高于Python循环:

# 只合并需要的州级字段,减少内存占用
merged_survey = pd.merge(
    Surveyframe,
    taxframe_wide[['Location', 'Year', 'Total_Tax/Pack', 'Avg_Cost/Pack']],
    on=['Location', 'Year'],
    how='left'  # 保留所有个体数据,匹配不到的州/年份字段会显示NaN
)

效率说明

Pandas的pivot和merge都是对整个DataFrame批量操作,避免了逐行循环的Python级开销。针对20万+个体数据,这种方法能在几秒内完成,而循环实现可能需要数分钟甚至更久。

内容的提问来源于stack exchange,提问作者SammyRichards

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 01:01:05