You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用llama-index的RagEvaluatorPack遇ValueError:空字符串无法转浮点数

使用Llama-index结合Ragas执行RagEvaluatorPack时的ValueError解决方法

问题重现

在使用Llama-index的RagEvaluatorPack结合Ragas进行RAG系统评估时,触发了ValueError,错误提示为could not convert string to float: ''。相关代码如下:

judge_llm = OpenAI(temperature=0, model="gpt-3.5-turbo")

RagEvaluatorPack = download_llama_pack("RagEvaluatorPack", "./pack")
rag_evaluator = RagEvaluatorPack(
    query_engine=query_engine,
    rag_dataset=rag_dataset,  # defined in 1A
    judge_llm=judge_llm,
    show_progress=True,
)

benchmark_df = await rag_evaluator.arun(
    batch_size=2,
    sleep_time_in_seconds=60,
)

报错详情

---------------------------------------------------------------------------
ValueError                                Traceback (most recent call last)
Cell In[23], line 30
     16 rag_evaluator = RagEvaluatorPack(
     17     query_engine=query_engine,
     18     rag_dataset=rag_dataset,  # defined in 1A
     19     judge_llm=judge_llm,
     20     show_progress=True,
     21 )
     23 ############################################################################
     24 # NOTE: If have a lower tier subscription for OpenAI API like Usage Tier 1 #
     25 # then you'll need to use different batch_size and sleep_time_in_seconds.  #
     26 # For Usage Tier 1, settings that seemed to work well were batch_size=5,   #
     27 # and sleep_time_in_seconds=15 (as of December 2023.)                      #
     28 ############################################################################
---> 30 benchmark_df = await rag_evaluator.arun(
     31     batch_size=2,  # batches the number of openai api calls to make
     32     sleep_time_in_seconds=60,  # seconds to sleep before making an api call
     33 )

File D:\documents\github\infinitejoy_courses\creating-gpt-chatbots-for-enterprise-useca-vt9QSr1Q-py3.10\lib\site-packages\llama_index\packs\rag_evaluator\base.py:442, in RagEvaluatorPack.arun(self, batch_size, sleep_time_in_seconds)
    440 # which is heavily rate-limited
    441 eval_batch_size = int(max(batch_size / 4, 1))
--> 442 return await self._amake_evaluations(
    443     batch_size=eval_batch_size, sleep_time_in_seconds=eval_sleep_time_in_seconds
    444 )

File D:\documents\github\infinitejoy_courses\creating-gpt-chatbots-for-enterprise-useca-vt9QSr1Q-py3.10\lib\site-packages\llama_index\packs\rag_evaluator\base.py:366, in RagEvaluatorPack._amake_evaluations(self, batch_size, sleep_time_in_seconds)
    364 # do this in batches to avoid RateLimitError
    365 try:
--> 366     eval_results: List[EvaluationResult] = await asyncio.gather(*tasks)
    367 except RateLimitError as err:
    368     if self.show_progress:

File D:\ProgramData\miniconda3\lib\asyncio\tasks.py:304, in Task.__wakeup(self, future)
    302 def __wakeup(self, future):
    303     try:
--> 304         future.result()
    305     except BaseException as exc:
    306         # This may also be a cancellation.
    307         self.__step(exc)

File D:\ProgramData\miniconda3\lib\asyncio\tasks.py:232, in Task.__step(***failed resolving arguments***)
    228 try:
    229     if exc is None:
    230         # We use the `send` method directly, because coroutines
    231         # don't have `__iter__` and `__next__` methods.
--> 232         result = coro.send(None)
    233     else:
    234         result = coro.throw(exc)

File D:\documents\github\infinitejoy_courses\creating-gpt-chatbots-for-enterprise-useca-vt9QSr1Q-py3.10\lib\site-packages\llama_index\core\evaluation\correctness.py:146, in CorrectnessEvaluator.aevaluate(***failed resolving arguments***)
    138 eval_response = await self._llm.apredict(
    139     prompt=self._eval_template,
    140     query=query,
    141     generated_answer=response,
    142     reference_answer=reference or "(NO REFERENCE ANSWER SUPPLIED)",
    143 )
    145 # Use the parser function
--> 146 score, reasoning = self.parser_function(eval_response)
    148 return EvaluationResult(
    149     query=query,
    150     response=response,
   (...)
    153     feedback=reasoning,
    154 )

File D:\documents\github\infinitejoy_courses\creating-gpt-chatbots-for-enterprise-useca-vt9QSr1Q-py3.10\lib\site-packages\llama_index\core\evaluation\eval_utils.py:183, in default_parser(eval_response)
    173 """
    174 Default parser function for evaluation response.
    175 
   (...)
    180     Tuple[float, str]: A tuple containing the score as a float and the reasoning as a string.
    181 """
    182 score_str, reasoning_str = eval_response.split("\n", 1)
--> 183 score = float(score_str)
    184 reasoning = reasoning_str.lstrip("\n")
    185 return score, reasoning

ValueError: could not convert string to float: ''

依赖配置

python = ">=3.10,<3.12"
streamlit = "^1.31.1"
llama-index = "^0.10.9"
llama-index-embeddings-huggingface = "^0.1.1"
llama-index-llms-ollama = "^0.1.1"
ragas = "^0.1.2"
spacy = "^3.7.4"

错误原因

错误出现在默认评估结果解析器default_parser中:该函数假设LLM返回的内容是“分数\n推理过程”的格式,当LLM返回的内容为空字符串,或者拆分后的第一个部分为空时,尝试将空字符串转为浮点数就会触发ValueError。常见诱因包括:

  • LLM返回的评估结果格式不符合预期(比如没有换行、分数缺失)
  • 参考答案为空导致LLM输出异常
  • 依赖版本兼容性问题

解决方法

1. 自定义鲁棒的解析函数

替换默认解析器,增加对空值、格式错误的容错处理:

def custom_parser(eval_response):
    eval_response = eval_response.strip()
    # 处理空响应
    if not eval_response:
        return 0.0, "LLM返回评估结果为空"
    
    # 尝试拆分分数和推理
    parts = eval_response.split("\n", 1)
    if len(parts) < 2:
        # 无换行时尝试提取开头的数字作为分数
        import re
        score_match = re.search(r'^\d+(\.\d+)?', eval_response)
        if score_match:
            score_str = score_match.group()
            reasoning = eval_response[len(score_str):].strip() or "无推理内容"
        else:
            return 0.0, f"无法提取有效分数,响应内容:{eval_response}"
    else:
        score_str, reasoning = parts
        score_str = score_str.strip()
        reasoning = reasoning.strip() or "无推理内容"
    
    # 尝试转换分数
    try:
        score = float(score_str)
        # 确保分数在0-1范围内
        score = max(0.0, min(1.0, score))
    except ValueError:
        return 0.0, f"分数解析失败,原始字符串:{score_str},推理内容:{reasoning}"
    
    return score, reasoning

将自定义解析器传入评估器,示例如下:

from llama_index.core.evaluation import CorrectnessEvaluator

# 创建带自定义解析器的评估器
correctness_evaluator = CorrectnessEvaluator(
    llm=judge_llm,
    parser_function=custom_parser
)
# 若RagEvaluatorPack支持传入自定义评估器,将其替换默认实例即可

2. 强制LLM输出符合要求的格式

修改评估用的prompt模板,明确要求LLM返回“分数+换行+推理”的格式,比如在prompt末尾添加:

请严格按照以下格式返回评估结果:首先输出一个0到1之间的浮点数作为正确性分数,然后换行详细说明你的推理过程。

3. 升级依赖版本

当前使用的llama-index 0.10.9和ragas 0.1.2存在兼容性或格式解析的bug,尝试升级到稳定新版本:

pip install --upgrade llama-index ragas

4. 检查并处理数据集

确保rag_dataset中的reference_answer字段不为空,或者在代码中对空参考答案做特殊处理,避免LLM因为缺少参考信息返回异常格式的内容。

内容的提问来源于stack exchange,提问作者joydeep bhattacharjee

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.29 01:19:58