使用llama-index的RagEvaluatorPack遇ValueError:空字符串无法转浮点数
使用Llama-index结合Ragas执行RagEvaluatorPack时的ValueError解决方法
问题重现
在使用Llama-index的RagEvaluatorPack结合Ragas进行RAG系统评估时,触发了ValueError,错误提示为could not convert string to float: ''。相关代码如下:
judge_llm = OpenAI(temperature=0, model="gpt-3.5-turbo") RagEvaluatorPack = download_llama_pack("RagEvaluatorPack", "./pack") rag_evaluator = RagEvaluatorPack( query_engine=query_engine, rag_dataset=rag_dataset, # defined in 1A judge_llm=judge_llm, show_progress=True, ) benchmark_df = await rag_evaluator.arun( batch_size=2, sleep_time_in_seconds=60, )
报错详情
--------------------------------------------------------------------------- ValueError Traceback (most recent call last) Cell In[23], line 30 16 rag_evaluator = RagEvaluatorPack( 17 query_engine=query_engine, 18 rag_dataset=rag_dataset, # defined in 1A 19 judge_llm=judge_llm, 20 show_progress=True, 21 ) 23 ############################################################################ 24 # NOTE: If have a lower tier subscription for OpenAI API like Usage Tier 1 # 25 # then you'll need to use different batch_size and sleep_time_in_seconds. # 26 # For Usage Tier 1, settings that seemed to work well were batch_size=5, # 27 # and sleep_time_in_seconds=15 (as of December 2023.) # 28 ############################################################################ ---> 30 benchmark_df = await rag_evaluator.arun( 31 batch_size=2, # batches the number of openai api calls to make 32 sleep_time_in_seconds=60, # seconds to sleep before making an api call 33 ) File D:\documents\github\infinitejoy_courses\creating-gpt-chatbots-for-enterprise-useca-vt9QSr1Q-py3.10\lib\site-packages\llama_index\packs\rag_evaluator\base.py:442, in RagEvaluatorPack.arun(self, batch_size, sleep_time_in_seconds) 440 # which is heavily rate-limited 441 eval_batch_size = int(max(batch_size / 4, 1)) --> 442 return await self._amake_evaluations( 443 batch_size=eval_batch_size, sleep_time_in_seconds=eval_sleep_time_in_seconds 444 ) File D:\documents\github\infinitejoy_courses\creating-gpt-chatbots-for-enterprise-useca-vt9QSr1Q-py3.10\lib\site-packages\llama_index\packs\rag_evaluator\base.py:366, in RagEvaluatorPack._amake_evaluations(self, batch_size, sleep_time_in_seconds) 364 # do this in batches to avoid RateLimitError 365 try: --> 366 eval_results: List[EvaluationResult] = await asyncio.gather(*tasks) 367 except RateLimitError as err: 368 if self.show_progress: File D:\ProgramData\miniconda3\lib\asyncio\tasks.py:304, in Task.__wakeup(self, future) 302 def __wakeup(self, future): 303 try: --> 304 future.result() 305 except BaseException as exc: 306 # This may also be a cancellation. 307 self.__step(exc) File D:\ProgramData\miniconda3\lib\asyncio\tasks.py:232, in Task.__step(***failed resolving arguments***) 228 try: 229 if exc is None: 230 # We use the `send` method directly, because coroutines 231 # don't have `__iter__` and `__next__` methods. --> 232 result = coro.send(None) 233 else: 234 result = coro.throw(exc) File D:\documents\github\infinitejoy_courses\creating-gpt-chatbots-for-enterprise-useca-vt9QSr1Q-py3.10\lib\site-packages\llama_index\core\evaluation\correctness.py:146, in CorrectnessEvaluator.aevaluate(***failed resolving arguments***) 138 eval_response = await self._llm.apredict( 139 prompt=self._eval_template, 140 query=query, 141 generated_answer=response, 142 reference_answer=reference or "(NO REFERENCE ANSWER SUPPLIED)", 143 ) 145 # Use the parser function --> 146 score, reasoning = self.parser_function(eval_response) 148 return EvaluationResult( 149 query=query, 150 response=response, (...) 153 feedback=reasoning, 154 ) File D:\documents\github\infinitejoy_courses\creating-gpt-chatbots-for-enterprise-useca-vt9QSr1Q-py3.10\lib\site-packages\llama_index\core\evaluation\eval_utils.py:183, in default_parser(eval_response) 173 """ 174 Default parser function for evaluation response. 175 (...) 180 Tuple[float, str]: A tuple containing the score as a float and the reasoning as a string. 181 """ 182 score_str, reasoning_str = eval_response.split("\n", 1) --> 183 score = float(score_str) 184 reasoning = reasoning_str.lstrip("\n") 185 return score, reasoning ValueError: could not convert string to float: ''
依赖配置
python = ">=3.10,<3.12" streamlit = "^1.31.1" llama-index = "^0.10.9" llama-index-embeddings-huggingface = "^0.1.1" llama-index-llms-ollama = "^0.1.1" ragas = "^0.1.2" spacy = "^3.7.4"
错误原因
错误出现在默认评估结果解析器default_parser中:该函数假设LLM返回的内容是“分数\n推理过程”的格式,当LLM返回的内容为空字符串,或者拆分后的第一个部分为空时,尝试将空字符串转为浮点数就会触发ValueError。常见诱因包括:
- LLM返回的评估结果格式不符合预期(比如没有换行、分数缺失)
- 参考答案为空导致LLM输出异常
- 依赖版本兼容性问题
解决方法
1. 自定义鲁棒的解析函数
替换默认解析器,增加对空值、格式错误的容错处理:
def custom_parser(eval_response): eval_response = eval_response.strip() # 处理空响应 if not eval_response: return 0.0, "LLM返回评估结果为空" # 尝试拆分分数和推理 parts = eval_response.split("\n", 1) if len(parts) < 2: # 无换行时尝试提取开头的数字作为分数 import re score_match = re.search(r'^\d+(\.\d+)?', eval_response) if score_match: score_str = score_match.group() reasoning = eval_response[len(score_str):].strip() or "无推理内容" else: return 0.0, f"无法提取有效分数,响应内容:{eval_response}" else: score_str, reasoning = parts score_str = score_str.strip() reasoning = reasoning.strip() or "无推理内容" # 尝试转换分数 try: score = float(score_str) # 确保分数在0-1范围内 score = max(0.0, min(1.0, score)) except ValueError: return 0.0, f"分数解析失败,原始字符串:{score_str},推理内容:{reasoning}" return score, reasoning
将自定义解析器传入评估器,示例如下:
from llama_index.core.evaluation import CorrectnessEvaluator # 创建带自定义解析器的评估器 correctness_evaluator = CorrectnessEvaluator( llm=judge_llm, parser_function=custom_parser ) # 若RagEvaluatorPack支持传入自定义评估器,将其替换默认实例即可
2. 强制LLM输出符合要求的格式
修改评估用的prompt模板,明确要求LLM返回“分数+换行+推理”的格式,比如在prompt末尾添加:
请严格按照以下格式返回评估结果:首先输出一个0到1之间的浮点数作为正确性分数,然后换行详细说明你的推理过程。
3. 升级依赖版本
当前使用的llama-index 0.10.9和ragas 0.1.2存在兼容性或格式解析的bug,尝试升级到稳定新版本:
pip install --upgrade llama-index ragas
4. 检查并处理数据集
确保rag_dataset中的reference_answer字段不为空,或者在代码中对空参考答案做特殊处理,避免LLM因为缺少参考信息返回异常格式的内容。
内容的提问来源于stack exchange,提问作者joydeep bhattacharjee
相关产品推荐
相关产品推荐

