Polars DataFrame字符串列前后追加字符时遇重复列名错误
解决Polars给所有字符串列同时添加前后缀的重复列名错误
问题场景
已有Polars DataFrame,能通过cs.string() + "||"给所有字符串列添加后缀,但执行"||" + cs.string() + "||"同时添加前后缀时,触发以下错误:
ComputeError: the name 'literal' passed to
LazyFrame.with_columnsis duplicate
It's possible that multiple expressions are returning the same default column name. If this is the case, try renaming the columns with.alias("new_name")to avoid duplicate column names.
错误原因
直接拼接字面量与列选择器时,Polars会将前后的"||"解析为默认列名为literal的临时列。当有多个字符串列时,每个列的处理都会生成名为literal的列,导致列名重复冲突。
解决方案
通过map方法对每个字符串列单独应用前后缀拼接逻辑,并保留原列名:
import polars as pl import polars.selectors as cs df = pl.DataFrame( { "A": [1, 2, 3], "B": ["apple", "orange", "grape"], "C": ["a", "b", "c"], } ) # 给所有字符串列前后添加"||" result = df.with_columns( cs.string().map(lambda col: pl.lit("||") + col + pl.lit("||")) ) print(result)
运行结果:
shape: (3, 3) ┌─────┬────────────┬───────┐ │ A ┆ B ┆ C │ │ --- ┆ --- ┆ --- │ │ i64 ┆ str ┆ str │ ╞═════╪════════════╪═══════╡ │ 1 ┆ ||apple|| ┆ ||a|| │ │ 2 ┆ ||orange|| ┆ ||b|| │ │ 3 ┆ ||grape|| ┆ ||c|| │ └─────┴────────────┴───────┘
另一种实现方式
也可以手动遍历字符串列,生成带别名的表达式列表:
string_cols = df.select(cs.string()).columns exprs = [pl.lit("||") + col + pl.lit("||").alias(col) for col in string_cols] result = df.with_columns(exprs)
这种方式更直观,明确指定每个处理后的列保留原名称,同样能避免列名重复问题。
内容的提问来源于stack exchange,提问作者blaylockbk
相关产品推荐
相关产品推荐

