You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

ADF管道调用Azure Function运行Python合并CSV无输出问题排查

问题修复方案

核心错误原因

  • 返回值类型不符合要求:你声明main函数返回func.HttpResponse类型,但实际返回普通字符串,会直接触发运行时InternalServerError
  • 写入目录无权限:Azure Function运行环境不允许直接在代码执行目录写入文件,你把CSV下载到当前工作目录会触发权限错误,本地能运行是因为本地目录有读写权限
  • 大概率缺少依赖声明:部署到云端的Python函数未配置requirements.txt,运行时找不到pandas、Azure存储相关依赖包也会触发报错
  • 可选网络访问限制:如果你的存储账户开启了防火墙,未将Function App的出站IP加入白名单也会导致存储读写失败,云端无报错但文件未生成

修复后的代码

init.py

import pandas as pd
import logging
from azure.storage.blob import BlobServiceClient
from azure.storage.filedatalake import DataLakeServiceClient
import azure.functions as func
import os

def main(req: func.HttpRequest) -> func.HttpResponse:
    logging.info('Python HTTP trigger function processed a request.')

    STORAGEACCOUNTURL= 'https://storage.blob.core.windows.net/'
    STORAGEACCOUNTKEY= '****'
    LOCALFILENAME= ['file1.csv', 'file2.csv']
    CONTAINERNAME= 'inputblob'
    # 改用Azure Function允许写入的临时目录
    TMP_PATH = '/tmp/'

    file1 = pd.DataFrame()
    file2 = pd.DataFrame()
    # 从blob下载
    blob_service_client_instance = BlobServiceClient(account_url=STORAGEACCOUNTURL, credential=STORAGEACCOUNTKEY)

    for i in LOCALFILENAME:
        file_path = os.path.join(TMP_PATH, i)
        with open(file_path, "wb") as my_blobs:
            blob_client_instance = blob_service_client_instance.get_blob_client(container=CONTAINERNAME, blob=i, snapshot=None)
            blob_data = blob_client_instance.download_blob()
            blob_data.readinto(my_blobs)
            if i == 'file1.csv':
                file1 = pd.read_csv(file_path)
            if i == 'file2.csv':
                file2 = pd.read_csv(file_path)
    
    # 合并文件
    summary = pd.merge(left=file1, right=file2, on='key', how='inner')
        
    # 上传到Data Lake
    service_client = DataLakeServiceClient(account_url="https://storage.dfs.core.windows.net/", credential='****')
    file_system_client = service_client.get_file_system_client(file_system="outputdatalake")
    directory_client = file_system_client.get_directory_client("functionapp") 
    file_client = directory_client.create_file("merged.csv") 
    file_contents = summary.to_csv(index=False)
    file_client.upload_data(file_contents, overwrite=True) 

    # 返回符合要求的HttpResponse对象
    return func.HttpResponse("This HTTP triggered function executed successfully.", status_code=200)

新增requirements.txt(放到项目根目录和function.json同层级)

pandas>=2.0.0
azure-storage-blob>=12.0.0
azure-storage-file-datalake>=12.0.0
azure-functions>=1.17.0

额外检查项

  • 存储账户如果开启了网络访问限制,需要将Function App的所有出站IP添加到存储账户的白名单
  • ADF调用Function时,Web活动的超时时间设置为大于函数实际执行时长,避免提前截断请求
  • Function App的运行时版本选择Python 3.8及以上,和本地开发版本保持一致

内容的提问来源于stack exchange,提问作者user17393771

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.26 23:06:03