You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

LLM评估中AnswerRelevancyMetric无结果且内核崩溃问题排查

问题解决方案

一、修复测试用例语法错误

你的LLMTestCase初始化代码存在语法错误,expected_output字段写法混乱,这是导致代码执行异常的直接原因之一。修正后的测试用例代码如下:

from deepeval.metrics import AnswerRelevancyMetric
from deepeval.test_case import LLMTestCase

# 初始化测试用例
test_case = LLMTestCase(
    input="The dog chased the cat up the tree, who ran up the tree?",
    actual_output="It depends, some might consider the cat, while others might argue the dog.",
    expected_output="The cat."
)

# 初始化指标
metric_ = AnswerRelevancyMetric(model=azure_openai)

二、解决metric_.score返回None的问题

metric_.score只有在调用metric_.measure(test_case)方法后才会被赋值,直接在measure()执行前打印必然返回None。正确执行顺序如下:

# 先执行评估逻辑
metric_.measure(test_case)
# 再打印评估分数
print(metric_.score)

三、修复Jupyter内核崩溃问题

1. 关闭流式输出

你在初始化AzureChatOpenAI时开启了streaming=True,同时自定义模型实现了异步方法a_generate,流式输出与异步调用的组合容易引发事件循环冲突,导致内核崩溃。将streaming改为False:

custom_model = AzureChatOpenAI(
    deployment_name=azure_openai_model, 
    api_key=azure_openai_secret.value, 
    azure_endpoint=azure_openai_endpoint, 
    api_version=azure_openai_api_version, 
    verbose=False, 
    streaming=False,  # 关闭流式输出
    temperature=0,
    callbacks=[StreamingStdOutCallbackHandler()]  # 若关闭流式,可移除该回调
)

2. 确保nest-asyncio正确应用

在代码最开头调用nest_asyncio.apply(),避免事件循环嵌套问题:

import nest_asyncio
nest_asyncio.apply()

# 后续自定义模型、测试用例代码放在此处

四、完整验证代码示例

整合所有修正后的完整代码:

import nest_asyncio
nest_asyncio.apply()

from langchain_openai import AzureChatOpenAI
from deepeval.models.base_model import DeepEvalBaseLLM
from langchain.callbacks.streaming_stdout import StreamingStdOutCallbackHandler

class AzureOpenAI(DeepEvalBaseLLM):
    def __init__(self, model):
        self.model = model

    def load_model(self):
        return self.model

    def generate(self, prompt: str) -> str:
        chat_model = self.load_model()
        return chat_model.invoke(prompt).content

    async def a_generate(self, prompt: str) -> str:
        chat_model = self.load_model()
        res = await chat_model.ainvoke(prompt)
        return res.content

    def get_model_name(self):
        return "Custom Azure OpenAI Model"

# 替换为你的实际参数
azure_openai_model = "your-deployment-name"
azure_openai_secret = "your-api-key"
azure_openai_endpoint = "your-azure-endpoint"
azure_openai_api_version = "your-api-version"

custom_model = AzureChatOpenAI(
    deployment_name=azure_openai_model, 
    api_key=azure_openai_secret, 
    azure_endpoint=azure_openai_endpoint, 
    api_version=azure_openai_api_version, 
    verbose=False, 
    streaming=False,
    temperature=0
)
azure_openai = AzureOpenAI(model=custom_model)

# 评估逻辑
from deepeval.metrics import AnswerRelevancyMetric
from deepeval.test_case import LLMTestCase

test_case = LLMTestCase(
    input="The dog chased the cat up the tree, who ran up the tree?",
    actual_output="It depends, some might consider the cat, while others might argue the dog.",
    expected_output="The cat."
)

metric_ = AnswerRelevancyMetric(model=azure_openai)
metric_.measure(test_case)
print(metric_.score)

内容的提问来源于stack exchange,提问作者Sara

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.20 18:40:04