You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python脚本命令行触发DuckDB析构错误,IPython执行正常求助

本地HuggingFace模型论文问答脚本的DuckDB析构错误解决

问题场景

使用本地HuggingFace模型实现PDF科研论文问答功能时:

  • IPython逐行执行:仅出现序列长度超限警告,可正常返回正确结果
  • 命令行执行python script.py:除相同警告外,脚本结束时触发DuckDB相关错误;即使修改最后一行打印结果,仍会在输出答案后出现错误

错误信息

Token indices sequence length is longer than the specified maximum sequence length for this model (1142 > 512). Running this sequence through the model will result in indexing errors
Exception ignored in: <function DuckDB.__del__ at 0x7f0d0b1b1fc0>
Traceback (most recent call last):
  File "/home/popsi/.local/lib/python3.10/site-packages/chromadb/db/duckdb.py", line 355, in __del__
AttributeError: 'NoneType' object has no attribute 'info'

原代码

import os
os.environ["HUGGINGFACEHUB_API_TOKEN"] = 'hf-xxxxxx'

from langchain.embeddings import HuggingFaceEmbeddings
from langchain import HuggingFaceHub
from langchain.llms import HuggingFacePipeline
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM, pipeline, AutoModelForSeq2SeqLM

model_id = 'google/flan-t5-large'
tokenizer = AutoTokenizer.from_pretrained(model_id,max_length=1500)

model = AutoModelForSeq2SeqLM.from_pretrained(model_id)

pipe = pipeline("text2text-generation", model=model, tokenizer=tokenizer, max_length=1500)
llm = HuggingFacePipeline(pipeline=pipe)

from langchain.document_loaders import UnstructuredPDFLoader
from langchain.indexes import VectorstoreIndexCreator
from langchain.text_splitter import CharacterTextSplitter

pdf_folder_path = "/mnt/d/test/langchain/pdfs"
loaders = [UnstructuredPDFLoader(os.path.join(pdf_folder_path, fn)) for fn in os.listdir(pdf_folder_path)]

index = VectorstoreIndexCreator(
   embedding=HuggingFaceEmbeddings(),
   text_splitter=CharacterTextSplitter(chunk_size=1000, chunk_overlap=0)).from_loaders(loaders)

from langchain.chains import RetrievalQA
chain = RetrievalQA.from_chain_type(llm=llm,
                                    chain_type="stuff",
                                    retriever=index.vectorstore.as_retriever(),
                                    input_key="question")

chain.run('What method was used in calculations?')

解决方案

1. 显式关闭DuckDB连接

在脚本末尾主动关闭向量库的数据库连接,避免析构时的资源异常:

# 在chain.run之后添加
index.vectorstore.db.close()

2. 指定Chromadb持久化目录

默认内存模式下,脚本结束时对象析构顺序混乱导致错误,指定持久化路径可让数据库生命周期管理更规范:

# 修改VectorstoreIndexCreator的参数,添加持久化配置
index = VectorstoreIndexCreator(
    embedding=HuggingFaceEmbeddings(),
    text_splitter=CharacterTextSplitter(chunk_size=1000, chunk_overlap=0),
    vectorstore_kwargs={"persist_directory": "./chroma_persist"}  # 自定义本地持久化目录
).from_loaders(loaders)

3. 手动触发垃圾回收

在脚本末尾主动删除相关对象并触发垃圾回收,确保资源正确释放:

import gc

# 在chain.run之后添加
del chain
del index
gc.collect()

附带解决序列长度警告

将文本拆分的chunk_size调整为模型支持的最大长度以内(flan-t5-large默认最大序列长度为512),避免警告:

text_splitter=CharacterTextSplitter(chunk_size=500, chunk_overlap=50)

内容的提问来源于stack exchange,提问作者Igor Popov

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.24 11:17:52