Polars DataFrame使用apply()处理多列触发TypeError错误排查
问题分析与解决
错误原因
你的代码存在两个核心问题:
- Polars
apply用法误解:df.apply(doc_check)是列级操作,会把每一列的Series(而非单行数据)传入函数,你误以为是逐行传递带列名的字典/对象,导致用字符串索引row['Stock_name']时触发类型错误。 - DataFrame修改方式错误:Polars是不可变数据结构,不能像Pandas那样直接用
df['doc_type_new'] = ...赋值添加新列,必须用with_columns方法。
修正方案
方案1:行级自定义函数(适合复杂逻辑)
通过pl.struct包装需要的列,实现逐行传递数据给自定义函数:
import polars as pl df = pl.DataFrame({ 'Stock_name': ['reserve', 'fullfilment', 'ntl', 'ntl'], 'doc_type': ['sales', 'moving', 'sales', 'corr'], }) def doc_check(row): if row['Stock_name'] in ('reserve', 'ntl'): return 'sales' if row['doc_type'] == 'sales' else row['doc_type'] return row['doc_type'] # 用struct包装列后逐行应用函数,添加新列 df = df.with_columns( pl.struct(['Stock_name', 'doc_type']).apply(doc_check).alias('doc_type_new') )
方案2:矢量化表达式(推荐,性能更优)
Polars原生支持矢量化表达式,无需逐行遍历,处理大数据时速度更快:
import polars as pl df = pl.DataFrame({ 'Stock_name': ['reserve', 'fullfilment', 'ntl', 'ntl'], 'doc_type': ['sales', 'moving', 'sales', 'corr'], }) df = df.with_columns( pl.when( pl.col('Stock_name').is_in(['reserve', 'ntl']) & (pl.col('doc_type') == 'sales') ).then('sales') .otherwise(pl.col('doc_type')) .alias('doc_type_new') )
执行后得到的结果:
shape: (4, 3) ┌─────────────┬──────────┬───────────────┐ │ Stock_name ┆ doc_type ┆ doc_type_new │ │ --- ┆ --- ┆ --- │ │ str ┆ str ┆ str │ ╞═════════════╪══════════╪═══════════════╡ │ reserve ┆ sales ┆ sales │ │ fullfilment ┆ moving ┆ moving │ │ ntl ┆ sales ┆ sales │ │ ntl ┆ corr ┆ corr │ └─────────────┴──────────┴───────────────┘
内容的提问来源于stack exchange,提问作者Masik
相关产品推荐
相关产品推荐

