You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何高效实现按年份分组将V1每行除以当年V2总和?

问题:按年份分组计算V1除以当年V2总和

首先是你的DataFrame构建代码:

import pandas as pd

date  = ['2015-02-03 23:00:00','2015-02-03 23:30:00','2016-02-04 00:00:00','2016-02-04 00:30:00']
v1 = [33.24  , 31.71  , 34.39  , 34.49 ]
v2 = [35.24  , 33.71  , 36.39  , 36.49 ]
    
df = pd.DataFrame({'V1':v1,'V2':v2}, index=pd.to_datetime(date))

输出结果:

V1     V2
2015-02-03 23:00:00  33.24  35.24
2015-02-03 23:30:00  31.71  33.71
2016-02-04 00:00:00  34.39  36.39
2016-02-04 00:30:00  34.49  36.49

需求:将V1列的每一行数据除以该年份下V2列的总和。

你的代码问题分析

你使用groupby(df.index.year).apply(lambda x: x["V1"]/x['V2'].sum())的方式,返回的是带有分组层级索引的Series,直接赋值给df["result"]会因为索引结构不匹配,导致无法和原DataFrame的行正确对齐,从而出现错误。

高效解决方案

推荐使用transform方法,它会将分组计算的聚合值(这里是每年V2的总和)自动广播到原DataFrame的每一行,保持索引一致,完美适配你的需求:

# 计算每年V2的总和,并广播到对应年份的每一行
yearly_v2_total = df.groupby(df.index.year)['V2'].transform('sum')
# 计算V1除以当年V2总和的结果
df['result'] = df['V1'] / yearly_v2_total

运行后得到的结果:

V1     V2    result
2015-02-03 23:00:00  33.24  35.24  0.485507
2015-02-03 23:30:00  31.71  33.71  0.461493
2016-02-04 00:00:00  34.39  36.39  0.478762
2016-02-04 00:30:00  34.49  36.49  0.480238

也可以用链式调用的写法更简洁:

df = df.assign(
    result=lambda x: x['V1'] / x.groupby(x.index.year)['V2'].transform('sum')
)

内容的提问来源于stack exchange,提问作者Peslier53

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.18 21:00:34