You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Langchain+本地GPT4All处理TXT转Neo4j时遇两类错误求助

Langchain+GPT4All导入文本到Neo4j的报错处理

第一个错误:KeyError: 'input_variables'

问题原因

创建PromptTemplate时使用了错误的参数variables,Langchain的PromptTemplate要求通过input_variables声明模板中的变量名,而非直接传入变量值。直接传递documents对象会导致框架找不到input_variables字段,触发KeyError。

解决方法

将variables参数替换为input_variables,值设为模板中用到的变量名列表(此处为["documents"]),之后通过format方法传入实际文档内容,或使用LLMChain绑定模板与模型。

第二个错误:上下文窗口超限(The prompt is 5161 tokens and the context window is 2048)

问题原因

GPT4All-falcon-q4_0.gguf的上下文窗口上限为2048 tokens,你设置的max_tokens是模型生成内容的最大token数,与输入上下文长度无关。即便做了文本分片,一次性传入所有分片仍会超出输入上限。

解决方法

需逐个处理每个文本分片,对每个chunk单独调用模型提取关系,避免一次性传入大量内容。同时要确保每个分片长度加上prompt模板长度不超过模型上下文窗口。

修改后的完整代码

# Script to convert a corpus of many text files into a neo4j graph

# Imports
import os
from langchain.document_loaders import TextLoader
from langchain.text_splitter import RecursiveCharacterTextSplitter
from langchain.llms.gpt4all import GPT4All
from langchain.prompts import PromptTemplate
from langchain.chains import LLMChain
from transformers import AutoTokenizer

# Load tokenizer
tokenizer = AutoTokenizer.from_pretrained("bert-base-uncased")

def bert_len(text):
    """Return the length of a text in BERT tokens."""
    tokens = tokenizer.encode(text)
    return len(tokens)

def get_files(path: str) -> list:
    """Return a list of all files in a directory, recursively."""
    files = []
    for file in os.listdir(path):
        file_path = os.path.join(path, file)
        if os.path.isdir(file_path):
            files.extend(get_files(file_path))
        else:
            files.append(file_path)
    return files

# Get the text files
all_txt_files = get_files('data')
raw_txt_files = []
for current_file in all_txt_files:
    raw_txt_files.extend(TextLoader(current_file, encoding='utf-8').load())

# Create a text splitter object that will help us split the text into chunks
# 调整chunk_size,确保加上prompt模板后不超过2048
text_splitter = RecursiveCharacterTextSplitter(
    chunk_size = 1500,  # 留足模板的token空间
    chunk_overlap = 128,
    length_function = bert_len,
    separators=['\n\n', '\n', ' ', ''],
)

# Split all text files into chunks
documents = []
for doc in raw_txt_files:
    documents.extend(text_splitter.create_documents([doc.page_content]))

# 正确创建PromptTemplate:声明input_variables
prompt_template = PromptTemplate(
    template = """
    You are a network graph maker who extracts terms and their relations from a given context.
    You are provided with a context chunk (delimited by ```). Your task is to extract the ontology
    of terms mentioned in the given context. These terms should represent the key concepts as per the context.
    
    Thought 1: While traversing through each sentence, Think about the key terms mentioned in it.
        Terms may include object, entity, location, organization, person,
        condition, acronym, documents, service, concept, etc.
        Terms should be as atomistic as possible
    
    Thought 2: Think about how these terms can have one on one relation with other terms.
        Terms that are mentioned in the same sentence or the same paragraph are typically related to each other.
        Terms can be related to many other terms
        
    Thought 3: Find out the relation between each such related pair of terms.

    Format your output as a list of json. Each element of the list contains
    a pair of terms and the relation between them, like the following:
    [
        {"node_1": "A concept from extracted ontology",
         "node_2": "A related concept from extracted ontology",
         "edge": "relationship between the two concepts, node_1 and node_2 in one or two sentences"},
        {"node_1": "A concept from extracted ontology",
         "node_2": "A related concept from extracted ontology",
         "edge": "relationship between the two concepts, node_1 and node_2 in one or two sentences"}
    ]
    Context Documents: ```{documents}```
    """,
    input_variables = ["documents"]  # 声明变量名
)

# Create a GPT4All object
llm = GPT4All(
    model=r"C:\Users\chalu\AppData\Local\nomic.ai\GPT4All\gpt4all-falcon-q4_0.gguf",
    n_threads=3,
    max_tokens=512,  # 生成内容的token数,不用超过模型上限
    verbose=True,
)

# 创建LLMChain绑定模板和模型
chain = LLMChain(llm=llm, prompt=prompt_template)

# 逐个处理每个文档分片
all_results = []
for idx, doc_chunk in enumerate(documents):
    print(f"Processing chunk {idx+1}/{len(documents)}")
    try:
        result = chain.run(documents=doc_chunk.page_content)
        all_results.append(result)
        print(f"Chunk {idx+1} processed successfully")
    except Exception as e:
        print(f"Error processing chunk {idx+1}: {str(e)}")

# 后续可以把all_results中的数据导入Neo4j
# 比如解析每个结果的JSON,然后用Neo4j Python driver创建节点和关系
print("All chunks processed. Results collected in all_results.")

额外说明

  1. 调整chunk_size为1500,因为prompt模板本身会占用数百个token,确保总输入不超过2048。
  2. 使用LLMChain管理模型与模板的调用,比直接调用llm(prompt)更规范,也方便后续扩展。
  3. 遍历所有文档分片逐个处理,避免一次性传入大量内容触发上下文超限。
  4. 模板中用```包裹documents,让模型更容易识别上下文边界。

内容的提问来源于stack exchange,提问作者wildcat89

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.06 00:02:06