FastAPI部署HuggingFace NER模型输出与本地运行不一致问题排查
问题排查与解决方案
核心问题分析
API与本地模型运行结果不一致,根源在于两个关键差异:
Transformers Pipeline参数不匹配
- main.py中初始化pipeline时使用
task="ner",未配置实体聚合策略和二进制输出参数; - b.py中使用
task="token-classification",并指定aggregation_strategy="average"和binary_output=True。
这两个参数直接影响实体合并逻辑与标签格式,导致API返回的实体被拆分成子词(如sudores拆为s/ud/ores),且标签带B-前缀,与本地结果不符。
- main.py中初始化pipeline时使用
不必要的URL编码
a.py中对文本做了URL编码,导致API收到编码后的字符串(如señor变为se%C3%B1or),模型处理错误输入自然输出异常。即使取消编码,Pipeline参数差异仍会导致结果不匹配。
具体修复步骤
1. 统一Pipeline初始化参数(修改main.py)
将main.py中的pipeline初始化代码与b.py保持完全一致:
pipe1 = pipeline( task="token-classification", model=model1.to("cpu"), tokenizer=tokenizer1, binary_output=True, aggregation_strategy="average" )
2. 移除不必要的URL编码(修改a.py)
直接传递原始文本作为JSON payload,requests会自动处理编码:
import requests text = "señor se presenta con fiebre y sudores fríos" # 直接使用原始文本,无需URL编码 payload = {"text": text} url = "http://localhost:8000/predict/" headers = {"Content-Type": "application/json"} response = requests.post(url, json=payload, headers=headers) print(response.status_code) print(response.json())
验证结果
修改后,API返回结果将与b.py本地运行结果完全一致:
{'predictions': [{'entity_group': 'DISO', 'score': 0.99900204, 'word': ' fiebre', 'start': 22, 'end': 28}, {'entity_group': 'DISO', 'score': 0.99870765, 'word': ' sudores fríos', 'start': 31, 'end': 44}]}
内容的提问来源于stack exchange,提问作者chancar
相关产品推荐
相关产品推荐

