You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Polars中处理自定义函数多行结果及文本多列多行解析?

Polars自定义函数解析文本为多列多行数据的解决方案

问题背景

现有如下Polars DataFrame,需将text列内容按逗号分割为多行,再将每行内容按空格拆分为word1、word2等多列:

import polars as pl
df = pl.DataFrame({
    'file': ['aaa.txt','bbb.txt'], 
    'text': ['my little pony, your big pony','apple+banana, cake+coke']
})

自定义函数myfunc单独测试可生成字典列表,但结合map_elements使用时触发报错:

def myfunc(p_str: str) -> list:
    res = []
    for line in p_str.split(','):
        x = line.strip().split(' ')
        res.append({f'word{e+1}': w for e, w in enumerate(x)})
    return res

# 报错代码
(df.with_columns(pl.struct(['text']).map_elements(lambda x: myfunc(x['text'])).alias('aaa'))
)

报错信息:

thread '' panicked at crates/polars-core/src/chunked_array/builder/list/anonymous.rs:161:69:
called Result::unwrap() on an Err value: InvalidOperation(ErrString("It is not possible to concatenate arrays of different data types."))
--- PyO3 is resuming a panic after fetching a PanicException from Python. ---

期望输出:

file     word1         word2   word3
aaa.txt  my            little  pony
aaa.txt  your          big     pony
bbb.txt  apple+banana
bbb.txt  cake+coke

报错原因

map_elements返回的字典列表中,每个字典的键数量不一致(部分有3个键,部分仅1个),Polars无法将这类异构数据自动转换为统一类型的列,导致类型匹配失败。

解决方案

优先使用Polars原生字符串处理函数,效率更高且避免类型问题,步骤如下:

  1. 按逗号分割text列,用explode展开为多行
  2. 去除每行前后空格后,按空格分割为列表列
  3. 将列表列拆分为多列并重命名为word1、word2...

完整代码

import polars as pl

df = pl.DataFrame({
    'file': ['aaa.txt','bbb.txt'], 
    'text': ['my little pony, your big pony','apple+banana, cake+coke']
})

result = (
    df
    # 按逗号分割text并展开为多行
    .with_columns(pl.col('text').str.split(',').alias('split_text'))
    .explode('split_text')
    # 去除空格后按空格分割为列表
    .with_columns(pl.col('split_text').str.strip().str.split(' ').alias('words'))
    # 拆分列表为多列并重命名
    .unnest('words')
    .rename({col: f"word{int(col.split('_')[1])+1}" for col in df.columns if col.startswith('word_')})
    # 移除中间临时列
    .drop('text', 'split_text')
    # 可选:将null替换为空字符串
    .fill_null('')
)

print(result)

输出结果

shape: (4, 4)
┌──────────┬──────────────┬─────────┬───────┐
│ file     ┆ word1        ┆ word2   ┆ word3 │
│ ---      ┆ ---          ┆ ---     ┆ ---   │
│ str      ┆ str          ┆ str     ┆ str   │
╞══════════╪══════════════╪═════════╪═══════╡
│ aaa.txt  ┆ my           ┆ little  ┆ pony  │
│ aaa.txt  ┆ your         ┆ big     ┆ pony  │
│ bbb.txt  ┆ apple+banana ┆         ┆       │
│ bbb.txt  ┆ cake+coke    ┆         ┆       │
└──────────┴──────────────┴─────────┴───────┘

自定义函数兼容方案(不推荐)

若必须使用自定义函数,需保证每个字典的键数量一致,不足的用空字符串填充:

def myfunc(p_str: str) -> list:
    res = []
    lines = [line.strip().split(' ') for line in p_str.split(',')]
    max_len = max(len(x) for x in lines) if lines else 0
    for x in lines:
        padded = x + [''] * (max_len - len(x))
        res.append({f'word{e+1}': w for e, w in enumerate(padded)})
    return res

result = (
    df
    .with_columns(pl.col('text').map_elements(myfunc).alias('aaa'))
    .explode('aaa')
    .unnest('aaa')
)

print(result)

此方法可得到正确结果,但性能远低于原生函数,仅适合小型数据集。


内容的提问来源于stack exchange,提问作者lmocsi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.28 10:05:16