You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Polars中如何使用group_by分组并应用自定义函数?

错误原因

Polars的map_groups()方法要求传入的自定义函数必须返回pl.DataFrame或pl.Series类型对象,你现有函数直接返回float类型的相关系数,不符合接口要求,因此抛出属性错误。

解决方案

方案1:修改自定义函数适配map_groups要求

仅需要将原函数的返回值包装为Polars Series即可:

import polars as pl
import pandas as pd
from scipy.stats import spearmanr

def get_score(df: pl.DataFrame) -> pl.Series:
    corr = spearmanr(df["prediction"], df["target"]).correlation
    # 将float结果包装为带字段名的Series
    return pl.Series("spearman_corr", [corr])

df = pd.DataFrame({
    "era": [1, 1, 1, 2, 2, 2, 5],
    "prediction": [2, 4, 5, 190, 1, 4, 1],
    "target": [1, 3, 2, 1, 43, 3, 1]
})
df_pl = pl.from_pandas(df)

# 加maintain_order=True保持和Pandas groupby一致的排序规则
correlations = df_pl.group_by("era", maintain_order=True).map_groups(get_score)

输出结果为Polars DataFrame,包含era和对应相关系数spearman_corr两列。

方案2:不修改原函数,使用agg+自定义函数实现

如果不想改动原有Pandas版本的get_score函数,可以用struct打包字段后直接应用函数:

# 沿用你原来的get_score函数
def get_score(df):
   return spearmanr(df["prediction"], df["target"]).correlation

correlations = df_pl.group_by("era", maintain_order=True).agg(
    pl.struct("prediction", "target")
    .apply(lambda x: get_score(x.to_pandas()))
    .alias("spearman_corr")
)

结果对齐说明

如果需要得到和你原Pandas代码完全一致的Series格式结果,可以在上述代码末尾追加转换逻辑:

# 转成Pandas Series,和原Pandas运行结果完全相同
correlations_pd = correlations.to_pandas().set_index("era")["spearman_corr"]

内容的提问来源于stack exchange,提问作者jbssm

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.30 10:24:03