ADF传递字符串格式Schema数组至Databricks引发IndexError求解
问题:ADF传递的Schema字符串无法在Databricks中转为数组导致索引错误
问题详情
- 从Azure Data Factory(ADF)的JSON表传递参数
schema_array,值为:"[('creationdate','timestamp'),('agent_name_txt','string'),('email','string'),('agent_hire_date','date'),('days_since_hire_text','int'),('department_auto','string')]" - Databricks Notebook接收后,该参数被识别为字符串类型(执行
print(type(schema_array))输出<class 'str'>),后续遍历索引时触发IndexError: string index out of range - 硬编码该值时功能正常,但尝试在ADF中直接设置为数组类型时,收到错误提示:
"The variable 'udf' of type 'String' cannot be initialized or updated with value of type 'Array'. The variable 'udf' only supports values of types 'String'"
解决方案
由于传递的字符串是Python列表(包含元组)的字符串表示形式,可使用Python标准库ast中的literal_eval()方法将其解析为真实的Python列表对象,这样就能正常遍历元组元素。
修改后的完整代码
import ast import pyspark.sql.types as t def string_to_datatype(datatype): if datatype == 'timestamp': final_datatype = t.TimestampType() elif datatype == 'integer' or datatype == 'int': final_datatype = t.IntegerType() elif datatype == 'date': final_datatype = t.DateType() else: final_datatype = t.StringType() return final_datatype # 将字符串解析为Python列表 schema_array = ast.literal_eval(schema_array) schema = t.StructType([]) for s in schema_array: schema = schema.add(t.StructField(s[0], string_to_datatype(s[1])))
内容的提问来源于stack exchange,提问作者TheBeast
相关产品推荐
相关产品推荐

