LLM评估中AnswerRelevancyMetric无结果且内核崩溃问题排查
问题解决方案
一、修复测试用例语法错误
你的LLMTestCase初始化代码存在语法错误,expected_output字段写法混乱,这是导致代码执行异常的直接原因之一。修正后的测试用例代码如下:
from deepeval.metrics import AnswerRelevancyMetric from deepeval.test_case import LLMTestCase # 初始化测试用例 test_case = LLMTestCase( input="The dog chased the cat up the tree, who ran up the tree?", actual_output="It depends, some might consider the cat, while others might argue the dog.", expected_output="The cat." ) # 初始化指标 metric_ = AnswerRelevancyMetric(model=azure_openai)
二、解决metric_.score返回None的问题
metric_.score只有在调用metric_.measure(test_case)方法后才会被赋值,直接在measure()执行前打印必然返回None。正确执行顺序如下:
# 先执行评估逻辑 metric_.measure(test_case) # 再打印评估分数 print(metric_.score)
三、修复Jupyter内核崩溃问题
1. 关闭流式输出
你在初始化AzureChatOpenAI时开启了streaming=True,同时自定义模型实现了异步方法a_generate,流式输出与异步调用的组合容易引发事件循环冲突,导致内核崩溃。将streaming改为False:
custom_model = AzureChatOpenAI( deployment_name=azure_openai_model, api_key=azure_openai_secret.value, azure_endpoint=azure_openai_endpoint, api_version=azure_openai_api_version, verbose=False, streaming=False, # 关闭流式输出 temperature=0, callbacks=[StreamingStdOutCallbackHandler()] # 若关闭流式,可移除该回调 )
2. 确保nest-asyncio正确应用
在代码最开头调用nest_asyncio.apply(),避免事件循环嵌套问题:
import nest_asyncio nest_asyncio.apply() # 后续自定义模型、测试用例代码放在此处
四、完整验证代码示例
整合所有修正后的完整代码:
import nest_asyncio nest_asyncio.apply() from langchain_openai import AzureChatOpenAI from deepeval.models.base_model import DeepEvalBaseLLM from langchain.callbacks.streaming_stdout import StreamingStdOutCallbackHandler class AzureOpenAI(DeepEvalBaseLLM): def __init__(self, model): self.model = model def load_model(self): return self.model def generate(self, prompt: str) -> str: chat_model = self.load_model() return chat_model.invoke(prompt).content async def a_generate(self, prompt: str) -> str: chat_model = self.load_model() res = await chat_model.ainvoke(prompt) return res.content def get_model_name(self): return "Custom Azure OpenAI Model" # 替换为你的实际参数 azure_openai_model = "your-deployment-name" azure_openai_secret = "your-api-key" azure_openai_endpoint = "your-azure-endpoint" azure_openai_api_version = "your-api-version" custom_model = AzureChatOpenAI( deployment_name=azure_openai_model, api_key=azure_openai_secret, azure_endpoint=azure_openai_endpoint, api_version=azure_openai_api_version, verbose=False, streaming=False, temperature=0 ) azure_openai = AzureOpenAI(model=custom_model) # 评估逻辑 from deepeval.metrics import AnswerRelevancyMetric from deepeval.test_case import LLMTestCase test_case = LLMTestCase( input="The dog chased the cat up the tree, who ran up the tree?", actual_output="It depends, some might consider the cat, while others might argue the dog.", expected_output="The cat." ) metric_ = AnswerRelevancyMetric(model=azure_openai) metric_.measure(test_case) print(metric_.score)
内容的提问来源于stack exchange,提问作者Sara
相关产品推荐
相关产品推荐

