使用Langchain查询Word文档报错及本地离线查询咨询
问题描述
尝试用Langchain批量查询Word文档时出现报错,Traceback如下:
Traceback (most recent call last): File C:\Program Files\Spyder\pkgs\spyder_kernels\py3compat.py:356 in compat_exec exec(code, globals, locals) File c:\data\langchain\langchaintest.py:44 index = VectorstoreIndexCreator().from_loaders(loaders) File ~\AppData\Roaming\Python\Python38\site-packages\langchain\indexes\vectorstore.py:72 in from_loaders docs.extend(loader.load()) File ~\AppData\Roaming\Python\Python38\site-packages\langchain\document_loaders\text.py:17 in load with open(self.file_path, encoding=self.encoding) as f: OSError: [Errno 22] Invalid argument:
注:报错信息中invalid argument: 后跟随的是Word文档的原始文本
使用的代码如下:
import os os.environ["OPENAI_API_KEY"] = "xxxxxx" import os import docx from langchain.document_loaders import TextLoader # Function to get text from a docx file def get_text_from_docx(file_path): doc = docx.Document(file_path) full_text = [] for paragraph in doc.paragraphs: full_text.append(paragraph.text) return '\n'.join(full_text) # Load multiple Word documents folder_path = 'C:/Data/langchain' word_files = [os.path.join(folder_path, file) for file in os.listdir(folder_path) if file.endswith('.docx')] loaders = [] for word_file in word_files: text = get_text_from_docx(word_file) loader = TextLoader(text) loaders.append(loader) from langchain.indexes import VectorstoreIndexCreator index = VectorstoreIndexCreator().from_loaders(loaders) query = "What are the main points discussed in the documents?" responses = index.query(query) print(responses) results_with_source=index.query_with_sources(query) print(results_with_source)
需要解决两个问题:
- 报错中
TextLoader需要传入的正确参数是什么? - 能否在无网络、不依赖OpenAI的本地环境实现此类查询?
解决方案
1. 修复TextLoader参数错误
报错原因是TextLoader的作用是加载本地文本文件,它的构造参数需要传入文件路径字符串,而你传入的是从Word文档中提取的原始文本内容,导致系统尝试把文本内容当作文件路径去打开,自然触发无效参数错误。
有两种修复方式:
方式一:直接使用Langchain的DocxLoader
Langchain已提供专门加载Word文档的DocxLoader,无需自定义文本提取函数,简化代码:
import os os.environ["OPENAI_API_KEY"] = "xxxxxx" from langchain.document_loaders import DocxLoader from langchain.indexes import VectorstoreIndexCreator folder_path = 'C:/Data/langchain' word_files = [os.path.join(folder_path, file) for file in os.listdir(folder_path) if file.endswith('.docx')] loaders = [DocxLoader(file) for file in word_files] index = VectorstoreIndexCreator().from_loaders(loaders) query = "What are the main points discussed in the documents?" responses = index.query(query) print(responses) results_with_source=index.query_with_sources(query) print(results_with_source)
方式二:保留自定义文本提取,改用Document对象
如果需要自定义Word文本提取逻辑,不要用TextLoader,直接构造Document对象后用VectorstoreIndexCreator的from_documents方法:
import os os.environ["OPENAI_API_KEY"] = "xxxxxx" import docx from langchain.docstore.document import Document from langchain.indexes import VectorstoreIndexCreator def get_text_from_docx(file_path): doc = docx.Document(file_path) full_text = [] for paragraph in doc.paragraphs: full_text.append(paragraph.text) return '\n'.join(full_text) folder_path = 'C:/Data/langchain' word_files = [os.path.join(folder_path, file) for file in os.listdir(folder_path) if file.endswith('.docx')] documents = [] for file in word_files: text = get_text_from_docx(file) doc = Document(page_content=text, metadata={"source": file}) documents.append(doc) index = VectorstoreIndexCreator().from_documents(documents) query = "What are the main points discussed in the documents?" responses = index.query(query) print(responses) results_with_source=index.query_with_sources(query) print(results_with_source)
2. 本地无网络、不依赖OpenAI的实现方案
可以实现,需替换Langchain中依赖OpenAI的组件为本地开源替代项,核心替换两部分:
- 嵌入模型:用
sentence-transformers系列(如all-MiniLM-L6-v2) - 大语言模型:用本地部署的开源模型(如Llama 2、Qwen、Mistral等),通过
CTransformers或HuggingFacePipeline加载
示例代码(以sentence-transformers为嵌入模型,本地Llama 2量化模型为例):
import os import docx from langchain.docstore.document import Document from langchain.indexes import VectorstoreIndexCreator # 本地嵌入模型 from langchain.embeddings import HuggingFaceEmbeddings # 本地LLM加载器 from langchain.llms import CTransformers # 初始化本地嵌入模型 embeddings = HuggingFaceEmbeddings(model_name="all-MiniLM-L6-v2") # 初始化本地LLM(假设已下载Llama 2量化模型文件llama-2-7b-chat.ggmlv3.q4_0.bin) llm = CTransformers( model="llama-2-7b-chat.ggmlv3.q4_0.bin", model_type="llama", config={"max_new_tokens": 256, "temperature": 0.01} ) # 加载Word文档并构造Document对象 def get_text_from_docx(file_path): doc = docx.Document(file_path) full_text = [] for paragraph in doc.paragraphs: full_text.append(paragraph.text) return '\n'.join(full_text) folder_path = 'C:/Data/langchain' word_files = [os.path.join(folder_path, file) for file in os.listdir(folder_path) if file.endswith('.docx')] documents = [] for file in word_files: text = get_text_from_docx(file) doc = Document(page_content=text, metadata={"source": file}) documents.append(doc) # 创建索引时指定本地嵌入模型和LLM index = VectorstoreIndexCreator( embedding=embeddings, llm=llm ).from_documents(documents) query = "What are the main points discussed in the documents?" responses = index.query(query) print(responses) results_with_source=index.query_with_sources(query) print(results_with_source)
注意:本地运行需提前下载对应模型文件,确保硬件性能足够(7B量化模型至少需8GB以上内存,GPU显存充足可提升速度)
内容的提问来源于stack exchange,提问作者bitterjam
相关产品推荐
相关产品推荐

