You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

RAGAS框架下Agent Goal Accuracy始终为0的问题排查求助

RAGAS Agent Goal Accuracy 始终为0的问题排查与解决

问题背景

我用ragas==0.2.15搭建了基于LangChain的投资研究助手,功能包括引导用户做出合理投资决策、通过yFinance和NewsAPI获取实时金融数据、用OpenAI Embeddings实现检索增强生成(RAG),并用RAGAS开展多轮评估。

测试场景:

输入查询:"现有三只股票——NVIDIA、Sofi technologies和Tesla,客户不太关注可持续性,但属于风险厌恶型,请推荐应购买哪只股票?"
助手响应:"我推荐NVIDIA给风险厌恶型客户。"

核心问题:用gpt-3.5-turbo评估React Agent系统时,无论是否提供参考值,AgentGoalAccuracy始终返回0分。

相关代码

if tool_usage:
    tool_calls = [ToolCall(name=tool, args={"company": company}) for tool in set(tool_usage)]
else:
    tool_calls = []
ragas_sample = MultiTurnSample(
    user_input=[
        RHMessage(content=message),
        RAMessage(content=reasoning_text, 
                tool_calls=tool_calls
                  )
    ],
    reference_tool_calls=tool_calls,
    reference_topics=["Stock price analysis and investment recommendations"],
    reference=answer
)


tool_score = ToolCallAccuracy().multi_turn_score(ragas_sample)
evaluator_llm = llm_factory(model="gpt-3.5-turbo")
goal_scorer = AgentGoalAccuracyWithReference(llm=evaluator_llm)

# 3️⃣ Run scoring in async context
async def run_goal_accuracy():
    agent_score = await goal_scorer.multi_turn_ascore(ragas_sample)
    print("Agent Goal Accuracy:", agent_score)  # score.score is a float from 0 to 1
    return agent_score

# 4️⃣ Run the async function
goal_score = asyncio.run(run_goal_accuracy())
print("====goalscore=====")
print(goal_score)
scorer = TopicAdherenceScore(llm = evaluator_llm, mode="recall")
async def evaluate_topic_adherence():
    score = await scorer.multi_turn_ascore(ragas_sample)
    return score

# Run and store the result
topic_score = asyncio.run(evaluate_topic_adherence())
print("Tool usage:", tool_usage)
print("ans usage:", answer)

ragas_scores = {
    "single_turn": single_scores,
    "multi_turn": {
        "ToolCallAccuracy": tool_score,
        "AgentGoalAccuracy": goal_score,
        "TopicAdherenceScore":float(topic_score) if topic_score is not None else None
        
    }
}

return answer, ragas_scores

可能的原因及修复方案

1. 参考答案设置逻辑错误

你当前将reference赋值为answer(即助手的响应),但AgentGoalAccuracyWithReference需要的是符合用户需求的标准答案,而非助手的输出。让评估器拿助手的回复和自身对比,逻辑完全倒置,必然导致0分。

修复:

  • 定义包含决策依据的明确参考答案:
    reference_answer = "推荐NVIDIA给风险厌恶型客户。NVIDIA营收稳定、市场份额领先,相比Sofi(金融科技赛道波动大)和Tesla(依赖政策与产能,股价波动剧烈),风险更低,更匹配风险厌恶型投资者的需求。"
    
  • 将ragas_sample中的reference=answer改为reference=reference_answer

2. MultiTurnSample 结构不符合要求

当前user_input同时包含用户消息和助手消息,不符合RAGAS多轮样本的格式规范。正确的多轮样本应将每一轮的用户提问与助手响应作为独立对话单元。

修复:
调整ragas_sample的构造方式:

ragas_sample = MultiTurnSample(
    conversation=[
        UserTurn(
            user_input=RHMessage(content=message),
            assistant_output=RAMessage(content=reasoning_text, tool_calls=tool_calls)
        )
    ],
    reference_tool_calls=reference_tool_calls,
    reference_topics=["Stock price analysis and investment recommendations"],
    reference=reference_answer
)

3. 参考答案缺少决策依据

AgentGoalAccuracyWithReference需要参考答案明确关联用户的核心需求(这里是"风险厌恶型"),仅给出结论无法让评估器判断助手的响应是否真正满足需求,会被判定为目标未达成。

修复:
确保参考答案包含需求匹配的逻辑说明,比如对比三只股票的风险特征,明确指出NVIDIA符合风险厌恶型需求的原因。

4. 工具调用参考值设置错误

你将reference_tool_calls设为助手实际调用的工具,但如果助手未调用必要的工具(比如未获取三只股票的波动率、营收稳定性数据),即使结论正确,评估器也会认为Agent未通过合理路径达成目标,从而给出0分。

修复:
定义应该调用的必要工具作为参考值:

reference_tool_calls = [
    ToolCall(name="yfinance", args={"company": "NVIDIA"}),
    ToolCall(name="yfinance", args={"company": "Sofi technologies"}),
    ToolCall(name="yfinance", args={"company": "Tesla"})
]

将ragas_sample中的reference_tool_calls替换为上述值,让评估器能判断Agent是否调用了支撑决策的必要工具。

内容的提问来源于stack exchange,提问作者Divya M

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.12 14:51:00