You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Polars DataFrame使用apply()处理多列触发TypeError错误排查

问题分析与解决

错误原因

你的代码存在两个核心问题:

  1. Polars apply用法误解:df.apply(doc_check)是列级操作,会把每一列的Series(而非单行数据)传入函数,你误以为是逐行传递带列名的字典/对象,导致用字符串索引row['Stock_name']时触发类型错误。
  2. DataFrame修改方式错误:Polars是不可变数据结构,不能像Pandas那样直接用df['doc_type_new'] = ...赋值添加新列,必须用with_columns方法。

修正方案

方案1:行级自定义函数(适合复杂逻辑)

通过pl.struct包装需要的列,实现逐行传递数据给自定义函数:

import polars as pl

df = pl.DataFrame({
    'Stock_name': ['reserve', 'fullfilment', 'ntl', 'ntl'],
    'doc_type': ['sales', 'moving', 'sales', 'corr'],
})

def doc_check(row):
    if row['Stock_name'] in ('reserve', 'ntl'):
        return 'sales' if row['doc_type'] == 'sales' else row['doc_type']
    return row['doc_type']

# 用struct包装列后逐行应用函数,添加新列
df = df.with_columns(
    pl.struct(['Stock_name', 'doc_type']).apply(doc_check).alias('doc_type_new')
)

方案2:矢量化表达式(推荐,性能更优)

Polars原生支持矢量化表达式,无需逐行遍历,处理大数据时速度更快:

import polars as pl

df = pl.DataFrame({
    'Stock_name': ['reserve', 'fullfilment', 'ntl', 'ntl'],
    'doc_type': ['sales', 'moving', 'sales', 'corr'],
})

df = df.with_columns(
    pl.when(
        pl.col('Stock_name').is_in(['reserve', 'ntl']) & (pl.col('doc_type') == 'sales')
    ).then('sales')
    .otherwise(pl.col('doc_type'))
    .alias('doc_type_new')
)

执行后得到的结果:

shape: (4, 3)
┌─────────────┬──────────┬───────────────┐
│ Stock_name  ┆ doc_type ┆ doc_type_new  │
│ ---         ┆ ---      ┆ ---           │
│ str         ┆ str      ┆ str           │
╞═════════════╪══════════╪═══════════════╡
│ reserve     ┆ sales    ┆ sales         │
│ fullfilment ┆ moving   ┆ moving        │
│ ntl         ┆ sales    ┆ sales         │
│ ntl         ┆ corr     ┆ corr          │
└─────────────┴──────────┴───────────────┘

内容的提问来源于stack exchange,提问作者Masik

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.23 18:51:06