如何在Polars中对混合数据类型DataFrame的字符串列执行strip操作?
在Polars中按列类型执行strip操作
需求:对DataFrame中的字符串列执行strip操作去除首尾空格,非字符串列不做处理。在Pandas中可通过遍历列并判断类型实现:
df_clean = df_select.copy() for col in df_select.columns: if df_select[col].dtype == 'object': df_clean[col] = df_select[col].str.strip()
以下是Polars中的实现方法:
首先构造示例DataFrame:
import polars as pl df = pl.DataFrame( { "ID": [1, 1, 1, 1,], "A": ["foo ", "ham", "spam ", "egg",], "L": ["A54", " A12", "B84", " C12"], } )
简洁实现方式
Polars中字符串列的类型为pl.Utf8,可直接选中所有该类型列并执行strip操作,无需手动遍历判断:
df_clean = df.with_columns( pl.col(pl.Utf8).str.strip() )
执行后df_clean的结果:
shape: (4, 3) ┌─────┬──────┬─────┐ │ ID ┆ A ┆ L │ │ --- ┆ --- ┆ --- │ │ i64 ┆ str ┆ str │ ╞═════╪══════╪═════╡ │ 1 ┆ foo ┆ A54 │ │ 1 ┆ ham ┆ A12 │ │ 1 ┆ spam ┆ B84 │ │ 1 ┆ egg ┆ C12 │ └─────┴──────┴─────┘
说明
pl.col(pl.Utf8)会筛选出所有字符串类型的列str.strip()对选中的列执行去除首尾空格的操作with_columns方法保留原DataFrame的所有列,仅修改目标列,非字符串列(如示例中的ID列)保持不变
内容的提问来源于stack exchange,提问作者Horseman
相关产品推荐
相关产品推荐

