Azure Function配置部署指导:ADLS多XML解析及创建报错问题
Azure Functions + ADF 处理ADLS XML文件全指南(含报错解决)
一、解决func new报错:Object reference not set to an instance of an object
- 更新Azure Functions Core Tools:终端执行
npm install -g azure-functions-core-tools@4 --unsafe-perm true,同时将VS Code的Azure Functions扩展更新到最新版本 - 重新初始化项目:删除当前可能损坏的项目文件夹,在空白文件夹打开VS Code,按F1执行
Azure Functions: Create New Project,通过向导选择Python、Blob Trigger模板并手动输入函数名,替代命令行创建 - 清理本地缓存:Windows删除
%LOCALAPPDATA%\AzureFunctionsTools文件夹;macOS/Linux删除~/.azurefunctions文件夹,重启VS Code
二、VS Code中配置Blob Trigger函数(适配ADLS Gen2)
1. 基础环境确认
- 安装Python 3.8/3.9/3.10(Azure Functions支持版本),在VS Code中选择对应解释器
- 把依赖写入
requirements.txt:azure-functions azure-storage-blob>=12.10.0 xmltodict pandas pyarrow
2. 配置Blob Trigger指向ADLS Gen2
- 修改生成的
function.json,指定监控的ADLS路径和连接名:{ "scriptFile": "__init__.py", "bindings": [ { "name": "myblob", "type": "blobTrigger", "direction": "in", "path": "你的解压容器名/xml文件存放文件夹/{name}", "connection": "ADLS_STORAGE_CONNECTION_STRING" } ] } - 在
local.settings.json中添加连接字符串(从Azure Portal存储账户访问密钥获取):{ "IsEncrypted": false, "Values": { "AzureWebJobsStorage": "你的函数运行存储连接字符串", "FUNCTIONS_WORKER_RUNTIME": "python", "ADLS_STORAGE_CONNECTION_STRING": "你的ADLS存储连接字符串" } }
3. 整合XML解析代码
- 替换
__init__.py默认代码,嵌入你的解析逻辑:import azure.functions as func from azure.storage.blob import BlobServiceClient import xmltodict import pandas as pd import io import os def main(myblob: func.InputStream): # 读取XML内容 xml_content = myblob.read().decode('utf-8') xml_data = xmltodict.parse(xml_content) # 替换为你的数据提取逻辑 extracted_data = [] for item in xml_data.get('root', {}).get('items', {}).get('item', []): extracted_data.append({ 'id': item.get('id'), 'name': item.get('name'), 'value': item.get('value') }) # 生成Parquet(替换为to_csv可生成CSV) df = pd.DataFrame(extracted_data) output_buffer = io.BytesIO() df.to_parquet(output_buffer, index=False) output_buffer.seek(0) # 写入目标ADLS容器 blob_service_client = BlobServiceClient.from_connection_string(os.environ['ADLS_STORAGE_CONNECTION_STRING']) target_container = blob_service_client.get_container_client('你的目标容器名') target_blob_name = f"解析结果/{myblob.name.split('/')[-1].replace('.xml', '.parquet')}" target_container.upload_blob(target_blob_name, output_buffer, overwrite=True)
三、Data Factory配置:触发函数的两种方式
方式1:Blob Trigger自动触发(推荐)
- 在ADF的解压/复制活动中,将解压后的XML文件写入到Function配置的ADLS容器文件夹
- 文件写入完成后,Blob Trigger会自动触发函数执行解析,无需额外ADF配置
方式2:ADF Web活动主动调用
- 发布Function到Azure后,从Azure Portal获取函数的HTTP触发URL和访问密钥
- 在ADF流水线中添加Web活动,配置:
- URL:函数触发URL(格式:
https://你的函数应用名.azurewebsites.net/api/ProcessXMLFunction?code=你的访问密钥) - 方法:POST
- 正文:可传递解压文件路径等参数(如需批量处理)
- URL:函数触发URL(格式:
- 将Web活动设置为解压/复制活动的后续步骤,配置成功依赖,确保解压完成后再调用函数
四、部署到Azure
- 在VS Code中按F1执行
Azure Functions: Deploy to Function App,选择提前在Azure Portal创建好的Python运行时Function App - 部署完成后,在Azure Portal的Function App应用设置中添加
ADLS_STORAGE_CONNECTION_STRING,替换本地配置
内容的提问来源于stack exchange,提问作者Mohammed Arif
相关产品推荐
相关产品推荐

