Pandas groupby触发FutureWarning的消除方案及目标输出实现问询
消除Pandas分组的FutureWarning且不添加分组列到DataFrame
我写了一段读取CSV的Python脚本,功能正常,但每次执行都会弹出这个FutureWarning:
FutureWarning: A grouping was used that is not in the columns of the DataFrame and so was excluded from the result. This grouping will be included in a future version of pandas. Add the grouping as a column of the DataFrame to silence this warning.
警告来自这段代码:
formatted_df = {k: 'first' for k in df.columns} | {'Commodity': tuple, 'FixedPriceStrike': tuple, 'Quantity': tuple} group = (df['TradeID'].str.strip().ne('') | df['TradeDate'].str.strip().ne('')).cumsum() df = df.groupby(group, as_index=False).agg(formatted_df)
示例数据
TradeID TradeDate Commodity StartDate ExpiryDate FixedPrice Quantity MTMValue -------- ---------- --------- --------- ---------- ---------- -------- --------- aaa 01/01/2024 commodity1 01/01/2024 01/01/2024 10.00 10 100.00 commodity2 10.00 10 bbb 01/01/2024 commodity1 01/01/2024 01/01/2024 10.00 10 100.00 commodity2 10.00 10 ccc 01/01/2024 commodity1 01/01/2024 01/01/2024 10.00 10 100.00 commodity2 10.00 10
预期输出
TradeID TradeDate Commodity StartDate ExpiryDate FixedPrice Quantity MTMValue -------- ---------- --------- --------- ---------- ---------- -------- --------- aaa 01/01/2024 (com1,com2) 01/01/2024 01/01/2024 (10,10) (10,10) 100.00 bbb 01/01/2024 (com1,com2) 01/01/2024 01/01/2024 (10,10) (10,10) 100.00 ccc 01/01/2024 (com1,com2) 01/01/2024 01/01/2024 (10,10) (10,10) 100.00
我不想把分组键添加到DataFrame里,求消除警告的方法或者实现预期输出的替代方案。
解决方法
方法1:临时屏蔽该FutureWarning
如果只是想消除警告、保留原有逻辑,可以用warnings模块精准屏蔽这个特定警告:
import warnings import pandas as pd # 屏蔽指定的FutureWarning warnings.filterwarnings("ignore", category=FutureWarning, message="A grouping was used that is not in the columns of the DataFrame and so was excluded from the result.") # 你的原有代码逻辑 formatted_df = {k: 'first' for k in df.columns} | {'Commodity': tuple, 'FixedPriceStrike': tuple, 'Quantity': tuple} group = (df['TradeID'].str.strip().ne('') | df['TradeDate'].str.strip().ne('')).cumsum() df = df.groupby(group, as_index=False).agg(formatted_df)
方法2:用transform+drop_duplicates替代分组聚合(无警告)
如果想从根源避免警告,可以换一种实现逻辑,不需要依赖外部分组键:
# 标记每个分组的行 group = (df['TradeID'].str.strip().ne('') | df['TradeDate'].str.strip().ne('')).cumsum() # 对需要合并成元组的列,用transform生成对应元组 cols_to_tuple = ['Commodity', 'FixedPriceStrike', 'Quantity'] for col in cols_to_tuple: df[col] = df.groupby(group)[col].transform(tuple) # 对其他列保留第一个非空值,然后去重得到最终结果 df = df.groupby(group, as_index=False).first()
这个逻辑和原代码效果完全一致,且不会触发警告,也无需将分组键保留在DataFrame中。
内容的提问来源于stack exchange,提问作者iBeMeltin
相关产品推荐
相关产品推荐

