You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Polars:结构体列表初始化DataFrame的类型异常及优化咨询

Polars结构体列初始化的类型问题与解决方案

一、为什么列表包裹结构体后类型变为List(Struct)?

Polars的DataFrame构造逻辑对不同输入类型的处理规则有差异:

  • 传入单个polars.struct对象时,Polars将其识别为单一行的结构体列,列类型为Struct。
  • 传入Python列表包裹的多个polars.struct对象时,Polars默认把这个列表解析为每个元素是List类型的值,而非多行结构体数据。这是因为Python列表本身是容器类型,Polars会将其映射为自身的List类型,最终列类型变成List(Struct)。

对比数值类型的情况:传入j=range(2)时,range是可迭代的数值序列,Polars会自动解析为多行数值列;但polars.struct是Polars专属对象,Python列表包裹它时,Polars不会自动将每个struct映射为单独的行。

二、更简洁的结构体列初始化方式

方法1:用polars.Series明确标记为结构体列

通过polars.Series包裹结构体列表,告诉Polars这是多行的结构体数据:

cols = list('ab')
df = polars.DataFrame(dict(
    j=polars.Series([
        polars.struct([polars.lit(j).alias(col) for j, col in enumerate(cols)], eager=True)
        for k in range(2)
    ])
))
print(df)
print(df.schema)

输出符合预期:

j
 {0,1}
 {0,1}
shape: (2, 1)
{'j': Struct([Field('a', Int32), Field('b', Int32)])}

方法2:用polars.repeat生成重复的结构体行

如果所有行的结构体内容相同,直接用repeat生成多行,无需Python循环:

cols = list('ab')
struct_val = polars.struct([polars.lit(j).alias(col) for j, col in enumerate(cols)], eager=True)
df = polars.DataFrame(dict(
    j=polars.repeat(struct_val, n=2, eager=True)
))

方法3:嵌套字典直接构造结构体列

如果结构体的每个字段对应多行数据,可直接传入嵌套字典,Polars会自动识别为结构体列:

df = polars.DataFrame({
    "j": {
        "a": [0, 0],
        "b": [1, 1]
    }
})

内容的提问来源于stack exchange,提问作者levant pied

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.21 03:28:13