You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Polars中按行判断B列字符串是否以A列字符串开头?

Polars实现按行判断B列字符串是否以A列开头

先看原始的DataFrame:

import polars as pl

df = pl.DataFrame({'A': ['a', 'b', 'c', 'd'], 'B': ['app', 'nop', 'cap', 'tab']})

输出结构:

shape: (4, 2)
┌─────┬─────┐
│ A   ┆ B   │
│ --- ┆ --- │
│ str ┆ str │
╞═════╪═════╡
│ a   ┆ app │
│ b   ┆ nop │
│ c   ┆ cap │
│ d   ┆ tab │
└─────┴─────┘

需求是新增列C,每行的C值为True当且仅当该行B列的字符串以A列的字符串开头,期望结果:

┌─────┬─────┬───────┐
│ A   ┆ B   ┆ C     │
│ --- ┆ --- ┆ ---   │
│ str ┆ str ┆ bool  │
╞═════╪═════╪═══════╡
│ a   ┆ app ┆ true  │
│ b   ┆ nop ┆ false │
│ c   ┆ cap ┆ true  │
│ d   ┆ tab ┆ false │
└─────┴─────┴───────┘

错误原因分析

你尝试的df['B'].str.starts_with(pl.col('A'))报错,是因为df['B']返回的是Polars Series,它的str.starts_with方法仅支持传入Python原生字符串/列表这类标量或可迭代对象,无法接收Polars的表达式对象(pl.col('A')是Expr类型)。

正确实现方式

方法一:使用Polars表达式(推荐,向量化性能更优)

通过pl.col构建列表达式,让str.starts_with接收另一个列的表达式,实现逐行判断:

df = df.with_columns(
    C=pl.col('B').str.starts_with(pl.col('A'))
)

这种方式是Polars的原生向量化操作,性能远高于逐行循环的方式。

方法二:逐行处理(适合复杂逻辑,性能稍差)

如果需要类似Pandasapply(axis=1)的逐行处理逻辑,可以用pl.row方法结合over实现:

df = df.with_columns(
    C=pl.row(lambda a, b: b.startswith(a), return_dtype=pl.Boolean).over(pl.int_range(0, pl.count()))
)

不过这种方式是逐行运算,数据量大时性能不如表达式方式,仅在复杂自定义逻辑场景下使用。

内容的提问来源于stack exchange,提问作者Syafiq Kamarul Azman

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 05:50:20