You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python-Polars聚合时如何计算众数?含字符串类型疑问

问题解决与优化方案

1. 修复pl.mode报错问题

你判断的原因完全正确,Polars没有全局的pl.mode()聚合函数,必须通过列表达式pl.col(col).mode()来调用众数聚合。直接替换expr_mode的生成逻辑即可:

expr_mode = [pl.col(col).mode().alias(f"mode_{col}") for col in cols]

注意:如果某列存在多个众数,mode()默认会返回所有众数组成的列表,你可以添加.first()来取第一个众数,比如pl.col(col).mode().first().alias(f"mode_{col}"),避免聚合后出现列表类型的列。

2. 适配字符串类型的聚合逻辑

当前代码仅针对数值类型(日期差转成的天数)处理,若要支持字符串列的聚合,需要区分列类型生成对应表达式:

  • 字符串列支持的聚合操作:max(字典序最大)、min(字典序最小)、mode(众数)
  • 字符串列不支持mean、std这类数值统计操作,强行调用会报错

修改后的代码如下,会自动区分列类型生成对应表达式:

def date_expr(df, df_base):
    # 改为在表达式中关联df_base的date_decision,避免提前join导致数据膨胀
    date_decision = pl.col("case_id").map_batches(
        lambda s: df_base.filter(pl.col("case_id").is_in(s))["date_decision"]
    )

    # 生成日期差表达式(仅处理结尾为"D"的日期列)
    date_cols = [col for col in df.columns if col[-1] == "D"]
    date_diff_exprs = [
        (pl.col(col) - date_decision).dt.total_days().alias(col)
        for col in date_cols
    ]

    # 分离数值列和字符串列(日期差转成数值,其他列可能为字符串)
    numeric_cols = date_cols
    string_cols = [col for col in df.columns if col not in date_cols and df[col].dtype == pl.Utf8]

    # 数值列聚合表达式
    numeric_exprs = []
    if numeric_cols:
        numeric_exprs.extend([pl.max(col).alias(f"max_{col}") for col in numeric_cols])
        numeric_exprs.extend([pl.min(col).alias(f"min_{col}") for col in numeric_cols])
        numeric_exprs.extend([pl.mean(col).alias(f"mean_{col}") for col in numeric_cols])
        numeric_exprs.extend([pl.col(col).mode().first().alias(f"mode_{col}") for col in numeric_cols])
        numeric_exprs.extend([pl.std(col).alias(f"std_{col}") for col in numeric_cols])

    # 字符串列聚合表达式
    string_exprs = []
    if string_cols:
        string_exprs.extend([pl.max(col).alias(f"max_{col}") for col in string_cols])
        string_exprs.extend([pl.min(col).alias(f"min_{col}") for col in string_cols])
        string_exprs.extend([pl.col(col).mode().first().alias(f"mode_{col}") for col in string_cols])

    return date_diff_exprs + numeric_exprs + string_exprs

# 调用逻辑
df = df.group_by("case_id").agg(date_expr(df, df_base))

3. 额外优化点

  • 避免提前全量join:原来的代码先对整个df做join会导致分组前数据膨胀,改用map_batches在聚合时关联df_base的date_decision,性能更优。
  • 合并列操作:将原来两次with_columns处理日期差的逻辑合并成一个表达式,代码更简洁。

内容的提问来源于stack exchange,提问作者Mihong Jelly

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.28 11:05:18