You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何批量去除Polars DataFrame中所有字符串的首尾空格?

解决Polars DataFrame去除字符串首尾空格的问题

你原来的代码没生效,是因为pl.all().map()里的lambda参数x是整个列对象,不是单个字符串元素,所以isinstance(x, str)判断永远不成立,自然不会执行strip()操作。

下面是两种正确的实现方式:

方法一:使用Polars内置字符串函数(推荐)

Polars提供了专门的字符串处理方法,属于向量化操作,效率远高于逐元素处理,能精准对所有字符串列生效:

import polars as pl

df = pl.DataFrame(
    {
        "A": ["foo ", "ham", "spam ", "egg"],
        "L": ["A54", " A12", "B84", " C12"],
        "Num": [1, 2, 3, 4]  # 测试非字符串列
    }
)

# 只清洗字符串类型的列,自动忽略非字符串列
df_clean_all = df.with_columns(pl.col(pl.String).str.strip())

print(df_clean_all)

输出结果:

shape: (4, 3)
┌──────┬──────┬─────┐
│ A    ┆ L    ┆ Num │
│ ---  ┆ ---  ┆ --- │
│ str  ┆ str  ┆ i64 │
╞══════╪══════╪═════╡
│ foo  ┆ A54  ┆ 1   │
│ ham  ┆ A12  ┆ 2   │
│ spam ┆ B84  ┆ 3   │
│ egg  ┆ C12  ┆ 4   │
└──────┴──────┴─────┘

方法二:逐元素处理(不推荐,仅作参考)

如果一定要沿用你原本的逐元素判断逻辑,需要用map_elements方法,它会遍历列中的每个元素:

df_clean = df.select(
    pl.all().map_elements(lambda x: x.strip() if isinstance(x, str) else x, return_dtype=pl.String)
)

注意需要指定return_dtype避免类型推断出错;这种方法效率低于内置函数,数据量大时不建议使用。

内容的提问来源于stack exchange,提问作者Horseman

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.02 07:10:28