You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在Foundry中使用*解包创建DataFrame时测试失败

问题:pytest夹具创建Spark DataFrame时的语法差异报错

我在pytest夹具中尝试用以下代码创建Spark DataFrame:

import pytest
from pyspark.sql import types as T

@pytest.fixture
def my_fun(spark_session):
    return spark_session.createDataFrame(
        [
            (*['test', 'testy'])
        ],
        T.StructType([
            T.StructField('mytest', T.StringType()),
            T.StructField('mytest2', T.StringType())
        ])
    )

def test_something(my_fun):
    return

运行后触发错误:

TypeError: StructType can not accept object 'test' in type <class 'str'>

换成('test', 'testy')替代(*['test', 'testy'])就能正常运行,为何二者不等价?(环境:Python 3.8.13,pytest-7.0.1)

原因分析

核心问题是两种写法生成的列表结构完全不同:

  • 使用(*['test', 'testy'])时,Python会把列表里的元素直接展开到外层列表中,最终生成的是['test', 'testy']——外层列表包含两个独立的字符串元素,而非一个包含两个元素的元组。
  • 而('test', 'testy')是完整的元组,作为外层列表的单个元素,最终生成[('test', 'testy')],完全符合Spark的要求。

Spark的createDataFrame要求传入的列表中,每个元素对应一行数据,必须是能匹配StructType结构的容器(元组、列表或Row对象)。当你传入单个字符串时,Spark会尝试把单个字符串映射到包含两个字段的StructType,自然触发类型不匹配的错误。

可以用简单打印验证这个差异:

print([(*['test', 'testy'])])  # 输出:['test', 'testy']
print([('test', 'testy')])     # 输出:[('test', 'testy')]

内容的提问来源于stack exchange,提问作者Benji

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 03:55:18