Python Flask应用中Haystack连续查询后出现内存错误的原因是什么?
问题:Haystack连续查询5次后触发内存错误
已在Elasticsearch中索引约1000份文档,使用Haystack执行查询时前几次可正常返回结果,但连续执行5次后出现内存错误导致程序终止。
复现代码
document_store = ElasticsearchDocumentStore(host="localhost", username="", password="", index="document") json_object = open("doc_json_file.json") data_json = json.load(json_object) json_object.close() document_store.write_documents(data_json) retriever = TfidfRetriever(document_store=document_store) reader = FARMReader(model_name_or_path="deepset/roberta-base-squad2", use_gpu=True) pipe = ExtractiveQAPipeline(reader, retriever) prediction = pipe.run(query=str(query), params={"Retriever": {"top_k": 20}, "Reader": {"top_k": 20}}) return prediction
错误日志
OSError: [WinError 1455] The paging file is too small for this operation to complete from .netcdf import netcdf_file, netcdf_variable File "<frozen importlib._bootstrap>", line 983, in _find_and_load File "<frozen importlib._bootstrap>", line 967, in _find_and_load_unlocked File "<frozen importlib._bootstrap>", line 677, in _load_unlocked File "<frozen importlib._bootstrap_external>", line 724, in exec_module File "<frozen importlib._bootstrap_external>", line 818, in get_code File "<frozen importlib._bootstrap_external>", line 917, in get_data MemoryError from pandas._libs.interval import Interval ImportError: DLL load failed: The paging file is too small for this operation to complete.
解决方法
- 复用核心实例,避免重复初始化:将
document_store、retriever、reader、pipe的初始化逻辑移到查询代码外部,仅执行一次,不要每次查询都重新创建实例,减少内存重复占用。 - 调整系统分页文件:Windows系统中,右键「此电脑」→ 属性 → 高级系统设置 → 高级 → 性能设置 → 高级 → 虚拟内存,勾选「自动管理所有驱动器的分页文件大小」,或手动增大分页文件的最小/最大值。
- 降低模型与查询负载:若GPU内存不足,将
use_gpu=True改为use_gpu=False,或换用轻量模型(如deepset/bert-base-squad2);同时降低top_k参数,比如把Retriever和Reader的top_k从20调至10,减少单次查询处理的文档数量。 - 主动释放内存:每次查询后调用
torch.cuda.empty_cache()(GPU场景)释放显存,或执行gc.collect()回收Python内存对象。 - 优化文档写入逻辑:确认文档已成功索引到Elasticsearch后,移除重复执行的
document_store.write_documents(data_json)操作,避免重复加载文档占用内存。
内容的提问来源于stack exchange,提问作者JAYAKUMAR S
相关产品推荐
相关产品推荐

