You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

ADF传递字符串格式Schema数组至Databricks引发IndexError求解

问题:ADF传递的Schema字符串无法在Databricks中转为数组导致索引错误

问题详情

  • 从Azure Data Factory(ADF)的JSON表传递参数schema_array,值为:
    "[('creationdate','timestamp'),('agent_name_txt','string'),('email','string'),('agent_hire_date','date'),('days_since_hire_text','int'),('department_auto','string')]"
    
  • Databricks Notebook接收后,该参数被识别为字符串类型(执行print(type(schema_array))输出<class 'str'>),后续遍历索引时触发IndexError: string index out of range
  • 硬编码该值时功能正常,但尝试在ADF中直接设置为数组类型时,收到错误提示:

    "The variable 'udf' of type 'String' cannot be initialized or updated with value of type 'Array'. The variable 'udf' only supports values of types 'String'"

解决方案

由于传递的字符串是Python列表(包含元组)的字符串表示形式,可使用Python标准库ast中的literal_eval()方法将其解析为真实的Python列表对象,这样就能正常遍历元组元素。

修改后的完整代码

import ast
import pyspark.sql.types as t

def string_to_datatype(datatype):
    if datatype == 'timestamp':
        final_datatype = t.TimestampType()
    elif datatype == 'integer' or datatype == 'int':
        final_datatype = t.IntegerType()
    elif datatype == 'date':
        final_datatype = t.DateType()
    else:
        final_datatype = t.StringType()
    return final_datatype

# 将字符串解析为Python列表
schema_array = ast.literal_eval(schema_array)

schema = t.StructType([])
for s in schema_array:
    schema = schema.add(t.StructField(s[0], string_to_datatype(s[1])))

内容的提问来源于stack exchange,提问作者TheBeast

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.02 15:05:17