能否用LogicApp获取SharePoint List Schema解决CSV数据类型问题?
一、使用Logic Apps获取SharePoint列表Schema
你可以通过两种方式在Logic Apps中获取列表的Schema信息:
方式1:借助触发器输出提取
- 在Azure门户创建或打开一个空白Logic App(Consumption/Standard计划均可)。
- 添加SharePoint触发器,比如当创建项时或获取项,配置目标SharePoint站点地址和列表名称,保存Logic App。
- 添加Compose动作,在输入框中引用触发器返回的列表项数据(例如
@{triggerBody()})。 - 手动触发Logic App一次,执行完成后查看Compose动作的输出,里面会包含带类型信息的列表字段结构,即Schema。
方式2:调用SharePoint REST API
- 在Logic Apps中添加HTTP动作,配置请求:
- 方法:
GET - URI:
https://你的SharePoint站点地址/_api/web/lists/getbytitle('你的列表名称')/fields - 请求头:添加
Accept: application/json;odata=verbose
- 方法:
- 运行Logic App后,HTTP动作的返回结果中,每个字段的
TypeAsString(字段类型)、Title(字段名)等信息就是完整的列表Schema。
二、将Schema应用到Databricks处理CSV
拿到Schema后,在Databricks中按以下步骤处理:
- 映射SharePoint字段类型到Spark支持的数据类型:比如SharePoint的
Number对应IntegerType/DoubleType,DateTime对应TimestampType,Boolean对应BooleanType,其余默认用StringType。 - 编写Spark代码读取CSV时指定自定义Schema,示例代码:
from pyspark.sql.types import StructType, StructField, IntegerType, DoubleType, TimestampType, BooleanType, StringType # 根据SharePoint Schema定义Spark Schema custom_schema = StructType([ StructField("ID", IntegerType(), nullable=True), StructField("产品名称", StringType(), nullable=True), StructField("价格", DoubleType(), nullable=True), StructField("发布日期", TimestampType(), nullable=True), StructField("是否可用", BooleanType(), nullable=True) ]) # 读取CSV并应用Schema df = spark.read.csv("/dbfs/路径/你的SharePoint导出文件.csv", header=True, schema=custom_schema)
- 处理完成后,可将DataFrame导出为BI工具兼容的格式,或直接让BI工具读取Databricks中的数据。
内容的提问来源于stack exchange,提问作者Dimitar Grigorov
相关产品推荐
相关产品推荐

