使用LLMGraphTransformer.convert_to_graph_documents()遇KeyError问题求助
问题描述
使用llm_transformer.convert_to_graph_documents()函数从PDF文本生成图文档并上传至Neo4j数据库时,偶尔会抛出KeyError:
KeyError: 'tail_type'
或
KeyError: 'tail'
该问题并非每次触发,通常在处理较长PDF时出现。相关代码如下:
# Extract text from PDF pdf_text = extract_text_from_pdf(pdf_path) # Create a Document object document_pdf = [Document(page_content=pdf_text)] llm_transformer = LLMGraphTransformer(llm=llm) # Create a GraphDocument object graph_documents = llm_transformer.convert_to_graph_documents(document_pdf)
错误原因
- LLM输出格式不稳定:
LLMGraphTransformer依赖大语言模型生成规范的三元组结构(头实体、关系、尾实体等),处理长文本时,模型可能因上下文长度限制或注意力分散,生成的结构化数据缺失tail或tail_type必填字段。 - 长文本未拆分:直接将整份PDF文本塞进单个
Document对象,当文本长度超出LLM上下文窗口时,模型输出会出现格式错乱,导致后续解析失败。 - 无格式校验逻辑:当前代码未对LLM生成的结果做合法性校验,一旦输出不符合预设结构,就会触发KeyError。
解决方法
- 拆分长文本为小文档:不要将整份PDF文本作为单个
Document,按页面或段落拆分,确保每个文档的长度在LLM上下文窗口范围内:# 示例:按页面拆分PDF文本(需实现按页提取的函数) pdf_pages = extract_text_from_pdf_by_pages(pdf_path) document_pdf = [Document(page_content=page) for page in pdf_pages] - 强制LLM输出规范格式:自定义
LLMGraphTransformer的prompt模板,明确要求必须包含所有必填字段,约束模型输出结构:from langchain_experimental.graph_transformers import LLMGraphTransformer custom_prompt = """ 请从给定文本中提取三元组,必须返回JSON数组格式,每个元素必须包含head、head_type、relation、tail、tail_type字段: [{"head": "实体1", "head_type": "实体类型", "relation": "关系", "tail": "实体2", "tail_type": "实体类型"}, ...] 文本:{text} """ llm_transformer = LLMGraphTransformer(llm=llm, prompt=custom_prompt) - 添加结果校验与容错:生成
graph_documents后,先过滤或修复无效三元组:graph_documents = llm_transformer.convert_to_graph_documents(document_pdf) # 校验并清理每个GraphDocument中的三元组 for doc in graph_documents: valid_triples = [] for triple in doc.triples: if all(key in triple for key in ["head", "head_type", "relation", "tail", "tail_type"]): valid_triples.append(triple) else: print(f"无效三元组已跳过:{triple}") doc.triples = valid_triples - 降低LLM输出随机性:如果模型支持温度参数,将温度值调低(比如设为0.1),减少输出的不确定性,提升格式稳定性。
内容的提问来源于stack exchange,提问作者ghinwamj
相关产品推荐
相关产品推荐

