You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Polars中group_by聚合调用collect_list报错:'Expr'对象无此属性

Polars group_by聚合列表报错及解决方法

问题描述

使用Polars进行group_by操作,目的是找出相同个人数据对应不同ClientID的分组,但运行代码时触发错误:Error: 'Expr' object has no attribute 'collect_list'。

原代码如下:

import polars as pl
from datetime import datetime

# Create a sample DataFrame with detailed personal data where all fields are the same except the client ID
df = pl.DataFrame({
    "PER_FirstName": ["John", "John", "John"],
    "PER_LastName": ["Doe", "Doe", "Doe"],
    "PER_DOB": [datetime(1990, 5, 1), datetime(1990, 5, 1), datetime(1990, 5, 1)],
    "PER_StreetAddress": ["123 Elm St", "123 Elm St", "123 Elm St"],
    "PER_ClientID": [101, 102, 103]
})


# Using group_by and trying a potentially available function.
try:
    result = df.group_by(['PER_FirstName', 'PER_LastName', 'PER_DOB', 'PER_StreetAddress']) \
               .agg([
                   pl.col('PER_ClientID').n_unique().alias('unique_client_ids'),
                   pl.col('PER_ClientID').collect_list().alias('client_ids')  # Trying collect_list
               ])
    print(result)
except AttributeError as e:
    print("Error:", e)

错误原因

你混淆了不同数据处理框架的API:collect_list是PySpark中的方法,Polars的Expr对象并没有这个属性。在Polars中,要聚合生成列值列表,需要使用专属的聚合函数或表达式方法。

修正后的代码

以下两种方法都可以解决问题,同时添加了过滤逻辑,直接筛选出存在多个不同ClientID的分组,更贴合你的需求:

方法1:使用pl.collect_list()聚合函数

import polars as pl
from datetime import datetime

df = pl.DataFrame({
    "PER_FirstName": ["John", "John", "John"],
    "PER_LastName": ["Doe", "Doe", "Doe"],
    "PER_DOB": [datetime(1990, 5, 1), datetime(1990, 5, 1), datetime(1990, 5, 1)],
    "PER_StreetAddress": ["123 Elm St", "123 Elm St", "123 Elm St"],
    "PER_ClientID": [101, 102, 103]
})

result = df.group_by(['PER_FirstName', 'PER_LastName', 'PER_DOB', 'PER_StreetAddress']) \
           .agg([
               pl.col('PER_ClientID').n_unique().alias('unique_client_ids'),
               pl.collect_list('PER_ClientID').alias('client_ids')
           ]) \
           .filter(pl.col('unique_client_ids') > 1)

print(result)

方法2:使用表达式的list()方法

import polars as pl
from datetime import datetime

df = pl.DataFrame({
    "PER_FirstName": ["John", "John", "John"],
    "PER_LastName": ["Doe", "Doe", "Doe"],
    "PER_DOB": [datetime(1990, 5, 1), datetime(1990, 5, 1), datetime(1990, 5, 1)],
    "PER_StreetAddress": ["123 Elm St", "123 Elm St", "123 Elm St"],
    "PER_ClientID": [101, 102, 103]
})

result = df.group_by(['PER_FirstName', 'PER_LastName', 'PER_DOB', 'PER_StreetAddress']) \
           .agg([
               pl.col('PER_ClientID').n_unique().alias('unique_client_ids'),
               pl.col('PER_ClientID').list().alias('client_ids')
           ]) \
           .filter(pl.col('unique_client_ids') > 1)

print(result)

输出结果

运行修正后的代码,会得到符合需求的分组:

shape: (1, 6)
┌────────────────┬───────────────┬─────────────────────┬──────────────────────┬──────────────────┬────────────┐
│ PER_FirstName  ┆ PER_LastName  ┆ PER_DOB             ┆ PER_StreetAddress    ┆ unique_client_ids ┆ client_ids │
│ ---            ┆ ---           ┆ ---                 ┆ ---                  ┆ ---              ┆ ---        │
│ str            ┆ str           ┆ datetime[μs]        ┆ str                  ┆ u32              ┆ list[i64]  │
╞════════════════╪═══════════════╪═════════════════════╪══════════════════════╪══════════════════╪════════════╡
│ John           ┆ Doe           ┆ 1990-05-01 00:00:00 ┆ 123 Elm St           ┆ 3                ┆ [101, 102, ┆
│                ┆               ┆                     ┆                      ┆                  ┆ 103]       │
└────────────────┴───────────────┴─────────────────────┴──────────────────────┴──────────────────┴────────────┘

内容的提问来源于stack exchange,提问作者tom

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.24 13:59:50