如何在Python中读取.doc文档流以实现LangChain自定义加载器?
扩展LangChain的CustomWordLoader以支持.doc格式流读取
我正在为LangChain开发CustomWordLoader,目前已通过python-docx实现了.docx文件的二进制流读取(文档流来自SharePoint站点),现有代码如下:
class CustomWordLoader(BaseLoader): """ This class is a custom loader for Word documents. It extends the BaseLoader class and overrides its methods. It uses the python-docx library to parse Word documents and optionally splits the text into manageable documents. Attributes: stream (io.BytesIO): A binary stream of the Word document. filename (str): The name of the Word document. """ def __init__(self, stream, filename: str): # Initialize with a binary stream and filename self.stream = stream self.filename = filename def load_and_split(self, text_splitter=None): # Use python-docx to parse the Word document from the binary stream doc = DocxDocument(self.stream) # Extract and concatenate all paragraph texts into a single string text = "\n".join([p.text for p in doc.paragraphs]) # Check if a text splitter utility is provided if text_splitter is not None: # Use the provided splitter to divide the text into manageable documents split_text = text_splitter.create_documents([text]) else: # Without a splitter, treat the entire text as one document split_text = [{'text': text, 'metadata': {'source': self.filename}}] # Add source metadata to each resulting document for doc in split_text: if isinstance(doc, dict): doc['metadata'] = {**doc.get('metadata', {}), 'source': self.filename} else: doc.metadata = {**doc.metadata, 'source': self.filename} return split_text
我的解决方案部署在基于3.11.8-alpine3.18的Docker容器中,出于安全考虑无法将文件下载到本地,需要像处理.docx那样直接读取二进制流,现寻找能读取.doc格式的替代方案。
可行的解决方案:使用textract库
1. 库选择理由
textract支持直接读取二进制流处理.doc文件,且能适配alpine容器环境(需安装对应的系统依赖),无需将文件保存到本地。
2. 修改后的CustomWordLoader代码
from langchain.document_loaders.base import BaseLoader from docx import Document as DocxDocument import textract import io class CustomWordLoader(BaseLoader): """ 自定义Word文档加载器,支持.docx和.doc格式的二进制流读取,可选择拆分文本。 属性: stream (io.BytesIO): Word文档的二进制流 filename (str): 文档文件名,用于判断格式 """ def __init__(self, stream: io.BytesIO, filename: str): self.stream = stream self.filename = filename def load_and_split(self, text_splitter=None): # 根据文件名后缀判断文档格式 if self.filename.lower().endswith('.docx'): doc = DocxDocument(self.stream) text = "\n".join([p.text for p in doc.paragraphs]) elif self.filename.lower().endswith('.doc'): # 读取.doc流并提取文本 text = textract.process(self.stream.read(), extension='doc').decode('utf-8') # 重置流指针,避免后续操作出错 self.stream.seek(0) else: raise ValueError(f"不支持的文件格式: {self.filename}") # 处理文本拆分逻辑 if text_splitter is not None: split_text = text_splitter.create_documents([text]) else: split_text = [{'text': text, 'metadata': {'source': self.filename}}] # 为每个文档添加来源元数据 for doc in split_text: if isinstance(doc, dict): doc['metadata'] = {**doc.get('metadata', {}), 'source': self.filename} else: doc.metadata = {**doc.metadata, 'source': self.filename} return split_text
3. Docker容器环境配置
由于textract处理.doc文件依赖antiword工具,需要在Dockerfile中添加系统依赖安装步骤:
FROM python:3.11.8-alpine3.18 # 安装antiword及编译依赖 RUN apk add --no-cache antiword python3-dev gcc musl-dev # 安装Python依赖 COPY requirements.txt . RUN pip install --no-cache-dir -r requirements.txt # 复制应用代码 COPY . /app WORKDIR /app CMD ["python", "your_main_script.py"]
requirements.txt需包含以下依赖:
langchain python-docx textract
内容的提问来源于stack exchange,提问作者Eric Vaillancourt
相关产品推荐
相关产品推荐

