You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Azure门户中Blob触发Python函数无法运行的排查求助

Azure Blob触发Python函数部署后不触发问题排查

问题现象

  • 本地VS Code开发的Blob触发Python函数,本地测试正常,部署到Azure函数应用后,向invoices容器上传Blob无触发,监控显示“无结果”。
  • 直接在Azure门户创建的Blob触发函数可正常触发。

相关代码

function.json

{
  "bindings": [
    {
      "name": "myblob",
      "path": "invoices",
      "connection": "my storage account's name",
      "direction": "in",
      "type": "blobTrigger"
    }
  ]
}

init.py

import logging
import azure.functions as func
from openpyxl import load_workbook
import openpyxl
import io
from azure.storage.blob import BlobServiceClient
import pandas as pd
from bs4 import BeautifulSoup

def main(myblob: func.InputStream):
    logging.info(f"Python blob trigger function processed blob 
"
                 f"Name: {myblob.name}
"
                 f"Blob Size: {myblob.length} bytes")
    conn_str1 = "DefaultEndpointsProtocol=https;AccountName=**********;AccountKey=*********;EndpointSuffix=core.windows.net"
    container1 = "invoices"
    blob_name = myblob.name[9:]
    blob_service_client1 = BlobServiceClient.from_connection_string(conn_str1)
    container_client1 = blob_service_client1.get_container_client(container1)
    download_blob1 = container_client1.download_blob(blob_name)
    soup = BeautifulSoup(download_blob1, features='xml')


    #Extracción de razon social, ruc y fecha de emision del xml
    rsocial = soup.find("cbc:RegistrationName").string.extract()
    print("Razón social = ", rsocial)
    ruc = soup.find("cac:AccountingSupplierParty").find("cbc:ID").string.extract()
    print("RUC = ", ruc)
    fecha = soup.find("cbc:IssueDate")
    print(" Fecha de emisión = ", fecha.string.extract())


    #Carga de tabla de Excel con el RUC de los proveedores de energía:
    conn_str2 = "DefaultEndpointsProtocol=https;AccountName=**********;AccountKey=*********;EndpointSuffix=core.windows.net"
    container2 = "files"
    xl_blob = "lista_ruc_2.xlsx"
    blob_service_client2 = BlobServiceClient.from_connection_string(conn_str2)
    container_client2 = blob_service_client2.get_container_client(container2)
    download_blob2 = container_client2.download_blob(xl_blob)
    file = io.BytesIO(download_blob2.readall())
    wb = openpyxl.load_workbook(file)  
    ws = wb["Tabelle1"]
    mapping = {}
    # Extracción del contenido de la tabla Excel y conversión a un pandas df
    for entry, data_boundary in ws.tables.items():
        data = ws[data_boundary]

        content = [[cell.value for cell in ent] 
                for ent in data
            ]
        
        header = content[0]
        rest = content[1:]
        
        df = pd.DataFrame(rest, columns = header)
        mapping[entry] = df

    #¿Es factura de energía? + Obtener n° de factura + Registro en el Excel de factuas correspondiente
    if int(ruc) in df['RUC'].unique():
        print('Es factura de energía')
        #Si es Luz del Sur:
        if int(ruc) == 20331898008:
            print('Es Luz del Sur')
            list_n_factura = soup.find_all("cbc:ID")
            n = 0
            while n <= len(list_n_factura):
                if n == 5:
                    n_factura = list_n_factura[n].string.extract()
                    print("N° factura = ", n_factura)
                n=n+1
        #Si es Enel:
        elif int(ruc) == 20269985900:
            print('Es Enel')
            list_n_factura = soup.find_all("cbc:ID")
            n = 0
            while n <= len(list_n_factura):
                if n == 1:
                    n_factura = list_n_factura[n].string.extract()
                    print("N° factura = ", n_factura)
                n=n+1
        #Si es Red de Energía del Perú:
        elif int(ruc) == 20504645046:
            print('Es Red de Energía del Perú')
            list_n_factura = soup.find_all("cbc:ID")
            n = 0
            while n <= len(list_n_factura):
                if n == 1:
                    n_factura = list_n_factura[n].string.extract()
                    print("N° factura = ", n_factura)
                n=n+1
        #Si es Agua Azul:
        elif int(ruc) == 20538865312:
            print('Es Agua Azul')
            list_n_factura = soup.find_all("cbc:ID")
            n = 0
            while n <= len(list_n_factura):
                if n == 0:
                    n_factura = list_n_factura[n].string.extract()
                    print("N° factura = ", n_factura)
                n=n+1
    else: 
        print('No es factura de energía')
        exit()

    #Post request con el diccionario conteniendo razon social y n° factura
    data = {'Razón Social': rsocial, 'N° Factura': n_factura}
    print(data)

    print('
------ FIN DEL PROGRAMA ------
')

问题排查与解决方案

1. 连接字符串配置错误

function.json中connection字段需填写Azure函数应用配置里的连接字符串名称(如AzureWebJobsStorage或自定义名称),而非存储账户名称。当前填写的"my storage account's name"不符合要求,需修正为正确的配置键名。

2. Blob触发器路径格式问题

path字段若要监听容器下所有Blob,应配置为invoices/{name},仅写invoices会导致触发器识别异常,无法正确监听Blob上传事件。

3. 依赖包缺失

本地测试环境安装的pandas、beautifulsoup4、openpyxl、azure-storage-blob等依赖,需在项目根目录创建requirements.txt文件并列出所有依赖包版本,部署时Azure会自动安装缺失的包,否则函数运行时会因依赖缺失崩溃,无触发记录。

4. 代码潜在错误

  • Blob名称截取逻辑:blob_name = myblob.name[9:]硬编码截取长度,若myblob.name格式变化会导致下载失败,改为blob_name = myblob.name.split('/')[-1]更可靠。
  • exit()终止函数:Azure函数中应正常返回或抛出异常,exit()会直接终止进程,被平台判定为异常,无触发日志记录,需替换为return或raise Exception()。
  • BeautifulSoup使用错误:string.extract()是冗余调用,直接取string即可(如rsocial = soup.find("cbc:RegistrationName").string),否则会引发AttributeError导致函数崩溃。

5. 存储账户权限问题

确保函数应用使用的身份(系统分配MSI/用户分配MSI)或连接字符串对应的账户密钥,对invoices容器有读取和列出权限,否则触发器无法获取Blob上传事件。

6. 触发器轮询延迟

Blob触发器依赖存储账户日志轮询,首次部署后可能有10-15分钟延迟,可尝试上传多个Blob等待,或检查存储账户“诊断设置”是否启用了日志记录。

内容的提问来源于stack exchange,提问作者Luis Chigne

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 11:44:56