Haystack PreProcessor报错:split_respect_sentence_boundary仅兼容split_by='word'
Haystack PreProcessor 报错问题解决
问题重现
用户代码(修正语法错误后):
from haystack.document_stores import InMemoryDocumentStore, SQLDocumentStore from haystack.nodes import TextConverter, PDFToTextConverter, PreProcessor from haystack.utils import clean_wiki_text, convert_files_to_docs, fetch_archive_from_http, print_answers doc_dir = "C:\\Users\\abcd\\Downloads\\PDF Files\\" docs = convert_files_to_docs(dir_path=doc_dir, clean_func=None, split_paragraphs=True) # 补充缺失的右括号 preprocessor = PreProcessor( clean_empty_lines=True, clean_whitespace=True, clean_header_footer=True, split_by="passage", split_length=2) doc = preprocessor.process(docs)
运行时抛出错误:
NotImplementedError: 'split_respect_sentence_boundary=True' is only compatible with split_by='word'
问题原因
Haystack的PreProcessor类中,split_respect_sentence_boundary参数的默认值为True,但该参数仅在split_by="word"时才被支持。当你设置split_by="passage"或split_by="sentence"时,默认开启的split_respect_sentence_boundary=True会触发不兼容错误,即使你没有显式声明这个参数。
解决方案
显式将split_respect_sentence_boundary设置为False,修改后的PreProcessor初始化代码如下:
preprocessor = PreProcessor( clean_empty_lines=True, clean_whitespace=True, clean_header_footer=True, split_by="passage", split_length=2, split_respect_sentence_boundary=False) # 新增该行
另外注意原代码中convert_files_to_docs一行末尾缺失右括号,需补充完整,否则会先触发语法错误。
内容的提问来源于stack exchange,提问作者atulbhushan
相关产品推荐
相关产品推荐

