You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Groq LLM与Llama Index在Python中生成失败的问题排查

LlamaParse + MarkdownElementNodeParser 400错误原因分析

问题代码

from llama_parse import LlamaParse
from llama_index.llms.groq import Groq
from llama_parse.base import ResultType, Language
from llama_index.core.node_parser import MarkdownElementNodeParser

parser = LlamaParse(
    api_key="xxx",
    result_type=ResultType.MD, 
    language=Language.ENGLISH, 
    parsing_instructions="""\
        The document is an exam documents. \
        It contains many tables and images. \
        Do not parse neither the definition, nor requirements, nor explanation sections.\
    """,
)

llm = Groq(
    model="llama3-8b-8192", 
    api_key="xxx",
)

documents = parser.load_data("path/to/file.pdf")

node_parser = MarkdownElementNodeParser(llm=llm, 
     summary_query_str="""
     The document is a property valuation documents. \
     It shows the key metrics for a property sale. \
     The paper includes detailed figures like the price, the size of the property, \
     the year of construction, the number of bedrooms and bathrooms. \
     Answer questions using the information in this document and be precise. \
     Skip lines for each bullet point. \
     Answer with only one number if asked for it.""")

nodes = node_parser.get_nodes_from_documents(documents)

触发错误

BadRequestError: Error code: 400 - {'error': {'message': "Failed to
call a function. Please adjust your prompt. See 'failed_generation'
for more details.", 'type': 'invalid_request_error', 'code':
'tool_use_failed', 'failed_generation': '\n{\n
"tool_calls": [\n {\n "id": "pending",\n "type":
"function",\n "function": {\n "name": "TableOutput"\n

},\n "parameters": {\n "table_id": "29 Jones St",\n

"table_title": "Property Information",\n "summary": "Summary of
property information"\n }\n }\n ]\n}\n'}}

错误原因

  • 任务描述严重冲突:LlamaParse的parsing_instructions声明要处理考试文档,跳过定义/要求/解释部分,但后续MarkdownElementNodeParser的summary_query_str又说这是房产估值文档。前后任务完全矛盾,导致LLM在处理表格时逻辑混乱,错误生成工具调用参数。
  • 工具调用参数非法:从failed_generation可以看到,LLM把房产地址"29 Jones St"当成table_id传入TableOutput工具,但table_id应该是系统内部的标识(比如数字),而非文本内容,这直接导致工具调用失败。
  • 节点解析指令偏离目标:summary_query_str的作用是指导LLM生成节点摘要,但你写的是问答式指令("Answer questions using the information..."),这会干扰LLM的工具调用逻辑,让它混淆了"生成节点"和"回答问题"的任务,进而触发错误的工具调用。

修正建议

  • 统一全流程任务描述:把LlamaParse的parsing_instructions改成房产估值相关的内容,比如:
    parsing_instructions="""\
        The document is a property valuation document. \
        Extract key metrics from tables, including price, property size, year of construction, number of bedrooms and bathrooms.\
    """
    
  • 修正summary_query_str的指令,聚焦节点摘要生成:
    summary_query_str="""
    Extract the key property metrics from this table: price, size, year of construction, number of bedrooms and bathrooms.
    Format the output as a clear bullet point list.
    """
    
  • 检查LlamaParse返回的Markdown格式,确保表格结构规范,没有乱码或格式错误。

内容的提问来源于stack exchange,提问作者Nicolas REY

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.20 07:17:09