You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Azure Function配置部署指导:ADLS多XML解析及创建报错问题

Azure Functions + ADF 处理ADLS XML文件全指南(含报错解决)

一、解决func new报错:Object reference not set to an instance of an object

  • 更新Azure Functions Core Tools:终端执行npm install -g azure-functions-core-tools@4 --unsafe-perm true,同时将VS Code的Azure Functions扩展更新到最新版本
  • 重新初始化项目:删除当前可能损坏的项目文件夹,在空白文件夹打开VS Code,按F1执行Azure Functions: Create New Project,通过向导选择Python、Blob Trigger模板并手动输入函数名,替代命令行创建
  • 清理本地缓存:Windows删除%LOCALAPPDATA%\AzureFunctionsTools文件夹;macOS/Linux删除~/.azurefunctions文件夹,重启VS Code

二、VS Code中配置Blob Trigger函数(适配ADLS Gen2)

1. 基础环境确认

  • 安装Python 3.8/3.9/3.10(Azure Functions支持版本),在VS Code中选择对应解释器
  • 把依赖写入requirements.txt:
    azure-functions
    azure-storage-blob>=12.10.0
    xmltodict
    pandas
    pyarrow
    

2. 配置Blob Trigger指向ADLS Gen2

  • 修改生成的function.json,指定监控的ADLS路径和连接名:
    {
      "scriptFile": "__init__.py",
      "bindings": [
        {
          "name": "myblob",
          "type": "blobTrigger",
          "direction": "in",
          "path": "你的解压容器名/xml文件存放文件夹/{name}",
          "connection": "ADLS_STORAGE_CONNECTION_STRING"
        }
      ]
    }
    
  • 在local.settings.json中添加连接字符串(从Azure Portal存储账户访问密钥获取):
    {
      "IsEncrypted": false,
      "Values": {
        "AzureWebJobsStorage": "你的函数运行存储连接字符串",
        "FUNCTIONS_WORKER_RUNTIME": "python",
        "ADLS_STORAGE_CONNECTION_STRING": "你的ADLS存储连接字符串"
      }
    }
    

3. 整合XML解析代码

  • 替换__init__.py默认代码,嵌入你的解析逻辑:
    import azure.functions as func
    from azure.storage.blob import BlobServiceClient
    import xmltodict
    import pandas as pd
    import io
    import os
    
    def main(myblob: func.InputStream):
        # 读取XML内容
        xml_content = myblob.read().decode('utf-8')
        xml_data = xmltodict.parse(xml_content)
        
        # 替换为你的数据提取逻辑
        extracted_data = []
        for item in xml_data.get('root', {}).get('items', {}).get('item', []):
            extracted_data.append({
                'id': item.get('id'),
                'name': item.get('name'),
                'value': item.get('value')
            })
        
        # 生成Parquet(替换为to_csv可生成CSV)
        df = pd.DataFrame(extracted_data)
        output_buffer = io.BytesIO()
        df.to_parquet(output_buffer, index=False)
        output_buffer.seek(0)
        
        # 写入目标ADLS容器
        blob_service_client = BlobServiceClient.from_connection_string(os.environ['ADLS_STORAGE_CONNECTION_STRING'])
        target_container = blob_service_client.get_container_client('你的目标容器名')
        target_blob_name = f"解析结果/{myblob.name.split('/')[-1].replace('.xml', '.parquet')}"
        target_container.upload_blob(target_blob_name, output_buffer, overwrite=True)
    

三、Data Factory配置:触发函数的两种方式

方式1:Blob Trigger自动触发(推荐)

  • 在ADF的解压/复制活动中,将解压后的XML文件写入到Function配置的ADLS容器文件夹
  • 文件写入完成后,Blob Trigger会自动触发函数执行解析,无需额外ADF配置

方式2:ADF Web活动主动调用

  • 发布Function到Azure后,从Azure Portal获取函数的HTTP触发URL和访问密钥
  • 在ADF流水线中添加Web活动,配置:
    • URL:函数触发URL(格式:https://你的函数应用名.azurewebsites.net/api/ProcessXMLFunction?code=你的访问密钥)
    • 方法:POST
    • 正文:可传递解压文件路径等参数(如需批量处理)
  • 将Web活动设置为解压/复制活动的后续步骤,配置成功依赖,确保解压完成后再调用函数

四、部署到Azure

  • 在VS Code中按F1执行Azure Functions: Deploy to Function App,选择提前在Azure Portal创建好的Python运行时Function App
  • 部署完成后,在Azure Portal的Function App应用设置中添加ADLS_STORAGE_CONNECTION_STRING,替换本地配置

内容的提问来源于stack exchange,提问作者Mohammed Arif

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.14 03:27:24