使用Groq LLM与Llama Index在Python中生成失败的问题排查
LlamaParse + MarkdownElementNodeParser 400错误原因分析
问题代码
from llama_parse import LlamaParse from llama_index.llms.groq import Groq from llama_parse.base import ResultType, Language from llama_index.core.node_parser import MarkdownElementNodeParser parser = LlamaParse( api_key="xxx", result_type=ResultType.MD, language=Language.ENGLISH, parsing_instructions="""\ The document is an exam documents. \ It contains many tables and images. \ Do not parse neither the definition, nor requirements, nor explanation sections.\ """, ) llm = Groq( model="llama3-8b-8192", api_key="xxx", ) documents = parser.load_data("path/to/file.pdf") node_parser = MarkdownElementNodeParser(llm=llm, summary_query_str=""" The document is a property valuation documents. \ It shows the key metrics for a property sale. \ The paper includes detailed figures like the price, the size of the property, \ the year of construction, the number of bedrooms and bathrooms. \ Answer questions using the information in this document and be precise. \ Skip lines for each bullet point. \ Answer with only one number if asked for it.""") nodes = node_parser.get_nodes_from_documents(documents)
触发错误
BadRequestError: Error code: 400 - {'error': {'message': "Failed to
call a function. Please adjust your prompt. See 'failed_generation'
for more details.", 'type': 'invalid_request_error', 'code':
'tool_use_failed', 'failed_generation': '\n{\n
"tool_calls": [\n {\n "id": "pending",\n "type":
"function",\n "function": {\n "name": "TableOutput"\n
},\n "parameters": {\n "table_id": "29 Jones St",\n
"table_title": "Property Information",\n "summary": "Summary of
property information"\n }\n }\n ]\n}\n'}}
错误原因
- 任务描述严重冲突:LlamaParse的
parsing_instructions声明要处理考试文档,跳过定义/要求/解释部分,但后续MarkdownElementNodeParser的summary_query_str又说这是房产估值文档。前后任务完全矛盾,导致LLM在处理表格时逻辑混乱,错误生成工具调用参数。 - 工具调用参数非法:从
failed_generation可以看到,LLM把房产地址"29 Jones St"当成table_id传入TableOutput工具,但table_id应该是系统内部的标识(比如数字),而非文本内容,这直接导致工具调用失败。 - 节点解析指令偏离目标:
summary_query_str的作用是指导LLM生成节点摘要,但你写的是问答式指令("Answer questions using the information..."),这会干扰LLM的工具调用逻辑,让它混淆了"生成节点"和"回答问题"的任务,进而触发错误的工具调用。
修正建议
- 统一全流程任务描述:把LlamaParse的
parsing_instructions改成房产估值相关的内容,比如:parsing_instructions="""\ The document is a property valuation document. \ Extract key metrics from tables, including price, property size, year of construction, number of bedrooms and bathrooms.\ """ - 修正
summary_query_str的指令,聚焦节点摘要生成:summary_query_str=""" Extract the key property metrics from this table: price, size, year of construction, number of bedrooms and bathrooms. Format the output as a clear bullet point list. """ - 检查LlamaParse返回的Markdown格式,确保表格结构规范,没有乱码或格式错误。
内容的提问来源于stack exchange,提问作者Nicolas REY
相关产品推荐
相关产品推荐

