Azure函数PDF上传+Gen AI分析问题求助:NoneType迭代错误
Azure函数开发问题:NoneType迭代错误与路由测试异常
问题描述
我是Azure函数新手,正在开发一个可上传PDF文档、通过生成式AI分析获取贷款风险洞察的函数,目前遇到以下问题:
- 触发
NoneType对象不可迭代错误 - 使用Postman测试时,GET方法无法上传文件
- 调用
/risk_prediction路由报错 /ingest路由无响应
错误信息
System.Private.CoreLib: Exception while executing function: Functions.risk_prediction. System.Private.CoreLib: Result: Failure
Exception: TypeError: 'NoneType' object is not iterable
代码实现
import logging import pathlib import tempfile import PyPDF2 import azure.functions as func from langchain_groq import ChatGroq app = func.FunctionApp(http_auth_level=func.AuthLevel.ANONYMOUS) groq_api_key=" key" @app.route(route="ingest", auth_level=func.AuthLevel.ANONYMOUS) def test(req: func.HttpRequest) -> func.HttpResponse: logging.info('Python HTTP trigger function processed a request.') try: # Fetch the file from the form-data file = req.files.file('file') # Fetch GUID from form-data guid = req.form.get('guid') if not guid and req.headers.get('Content-Type') == 'application/json': try: req_body = req.get_json() except ValueError: pass else: guid = req_body.get('guid') if not guid: return func.HttpResponse( "GUID is missing in the request.", status_code=400 ) suffix = pathlib.Path(file.filename).suffix with tempfile.NamedTemporaryFile(delete=False, suffix=suffix) as temp: temp.write(file.read()) temp_path = temp.name documents = [] with open(temp_path, 'rb') as file: reader = PyPDF2.PdfReader(file) text = '' for page_num in range(len(reader.pages)): page = reader.pages[page_num] text += page.extract_text() documents.append(text) message = f"data: {documents}" # Join all document texts into a single string return func.HttpResponse(message) except Exception as e: return func.HttpResponse(f"Info: Error Encountered. -- {e}", status_code=500) @app.route(route="risk_prediction") def risk_prediction(req: func.HttpRequest) -> func.HttpResponse: logging.info('Python HTTP trigger function processed a request.') documents=test(req) llm = ChatGroq(groq_api_key=groq_api_key, model_name="llama3-70b-8192") result='' prompt='''Provide valuable insight on loan risk for applicants, by anlyzing the documents {documents}''' for doc in documents: messages = [ ( "system",prompt, ), ("human", f'''contract document is {doc} and anlysis of previous document is {result}, include previous findings with current data and update response''')] result =llm.invoke(messages).content return result
测试URL
http://localhost:7071/api/testhttp://localhost:7071/api/risk_prediction
问题分析与修复方案
1. NoneType不可迭代错误根源
risk_prediction函数直接调用test(req)获取documents,但test返回的是func.HttpResponse对象(不是PDF文本列表)。若test返回错误响应,documents会成为无法迭代的对象,触发报错。
2. 核心修复点
(1)拆分PDF解析逻辑为独立函数
将文件解析与HTTP响应逻辑分离,避免路由函数耦合:
def parse_pdf(file) -> list: """解析PDF文件并返回文本列表""" documents = [] suffix = pathlib.Path(file.filename).suffix with tempfile.NamedTemporaryFile(delete=False, suffix=suffix) as temp: temp.write(file.read()) temp_path = temp.name with open(temp_path, 'rb') as f: reader = PyPDF2.PdfReader(f) text = '' for page_num in range(len(reader.pages)): page = reader.pages[page_num] text += page.extract_text() documents.append(text) # 清理临时文件 pathlib.Path(temp_path).unlink() return documents
(2)修正risk_prediction函数逻辑
不再调用路由函数获取数据,直接从请求中解析文件并分析:
@app.route(route="risk_prediction", auth_level=func.AuthLevel.ANONYMOUS, methods=['POST']) def risk_prediction(req: func.HttpRequest) -> func.HttpResponse: logging.info('Python HTTP trigger function processed a request.') try: file = req.files.get('file') if not file: return func.HttpResponse("PDF文件缺失", status_code=400) documents = parse_pdf(file) llm = ChatGroq(groq_api_key=groq_api_key, model_name="llama3-70b-8192") result = '' prompt = "基于提供的文档,分析贷款申请人的风险并给出有价值的洞察" for doc in documents: messages = [ ("system", prompt), ("human", f"合同文档内容:{doc}\n之前的分析结果:{result}\n结合已有内容更新风险分析") ] result = llm.invoke(messages).content return func.HttpResponse(result, status_code=200) except Exception as e: return func.HttpResponse(f"分析出错:{str(e)}", status_code=500)
(3)修复ingest路由的文件获取逻辑
原代码req.files.file('file')写法错误,改为req.files.get('file'),并明确支持POST方法:
@app.route(route="ingest", auth_level=func.AuthLevel.ANONYMOUS, methods=['POST']) def ingest(req: func.HttpRequest) -> func.HttpResponse: logging.info('Python HTTP trigger function processed a request.') try: file = req.files.get('file') if not file: return func.HttpResponse("PDF文件缺失", status_code=400) guid = req.form.get('guid') if not guid and req.headers.get('Content-Type') == 'application/json': try: req_body = req.get_json() except ValueError: pass else: guid = req_body.get('guid') if not guid: return func.HttpResponse("请求中缺少GUID", status_code=400) documents = parse_pdf(file) # 此处可添加文档与GUID关联存储的逻辑(如Azure Blob/Cosmos DB) return func.HttpResponse(f"文件已解析,GUID:{guid},内容长度:{len(documents[0])}", status_code=200) except Exception as e: return func.HttpResponse(f"处理出错:{str(e)}", status_code=500)
(4)Postman测试规范
- 必须使用POST方法,GET方法无法携带文件
- 请求头
Content-Type设置为multipart/form-data - 在
form-data中添加file字段(类型选文件)和guid字段(类型选文本)
3. 其他优化点
- 添加临时文件删除逻辑,避免磁盘占用
- 修正拼写错误:
anlyzing→analyzing,anlysis→analysis
内容的提问来源于stack exchange,提问作者prateek s
相关产品推荐
相关产品推荐

