You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中读取.doc文档流以实现LangChain自定义加载器?

扩展LangChain的CustomWordLoader以支持.doc格式流读取

我正在为LangChain开发CustomWordLoader,目前已通过python-docx实现了.docx文件的二进制流读取(文档流来自SharePoint站点),现有代码如下:

class CustomWordLoader(BaseLoader):
    """
    This class is a custom loader for Word documents. It extends the BaseLoader class and overrides its methods.
    It uses the python-docx library to parse Word documents and optionally splits the text into manageable documents.
    
    Attributes:
    stream (io.BytesIO): A binary stream of the Word document.
    filename (str): The name of the Word document.
    """
    def __init__(self, stream, filename: str):
        # Initialize with a binary stream and filename
        self.stream = stream
        self.filename = filename

    def load_and_split(self, text_splitter=None):
        # Use python-docx to parse the Word document from the binary stream
        doc = DocxDocument(self.stream)
        # Extract and concatenate all paragraph texts into a single string
        text = "\n".join([p.text for p in doc.paragraphs])

        # Check if a text splitter utility is provided
        if text_splitter is not None:
            # Use the provided splitter to divide the text into manageable documents
            split_text = text_splitter.create_documents([text])
        else:
            # Without a splitter, treat the entire text as one document
            split_text = [{'text': text, 'metadata': {'source': self.filename}}]

        # Add source metadata to each resulting document
        for doc in split_text:
            if isinstance(doc, dict):
                doc['metadata'] = {**doc.get('metadata', {}), 'source': self.filename}
            else:
                doc.metadata = {**doc.metadata, 'source': self.filename}

        return split_text

我的解决方案部署在基于3.11.8-alpine3.18的Docker容器中,出于安全考虑无法将文件下载到本地,需要像处理.docx那样直接读取二进制流,现寻找能读取.doc格式的替代方案。


可行的解决方案:使用textract库

1. 库选择理由

textract支持直接读取二进制流处理.doc文件,且能适配alpine容器环境(需安装对应的系统依赖),无需将文件保存到本地。

2. 修改后的CustomWordLoader代码

from langchain.document_loaders.base import BaseLoader
from docx import Document as DocxDocument
import textract
import io

class CustomWordLoader(BaseLoader):
    """
    自定义Word文档加载器,支持.docx和.doc格式的二进制流读取,可选择拆分文本。
    
    属性:
    stream (io.BytesIO): Word文档的二进制流
    filename (str): 文档文件名,用于判断格式
    """
    def __init__(self, stream: io.BytesIO, filename: str):
        self.stream = stream
        self.filename = filename

    def load_and_split(self, text_splitter=None):
        # 根据文件名后缀判断文档格式
        if self.filename.lower().endswith('.docx'):
            doc = DocxDocument(self.stream)
            text = "\n".join([p.text for p in doc.paragraphs])
        elif self.filename.lower().endswith('.doc'):
            # 读取.doc流并提取文本
            text = textract.process(self.stream.read(), extension='doc').decode('utf-8')
            # 重置流指针,避免后续操作出错
            self.stream.seek(0)
        else:
            raise ValueError(f"不支持的文件格式: {self.filename}")

        # 处理文本拆分逻辑
        if text_splitter is not None:
            split_text = text_splitter.create_documents([text])
        else:
            split_text = [{'text': text, 'metadata': {'source': self.filename}}]

        # 为每个文档添加来源元数据
        for doc in split_text:
            if isinstance(doc, dict):
                doc['metadata'] = {**doc.get('metadata', {}), 'source': self.filename}
            else:
                doc.metadata = {**doc.metadata, 'source': self.filename}

        return split_text

3. Docker容器环境配置

由于textract处理.doc文件依赖antiword工具,需要在Dockerfile中添加系统依赖安装步骤:

FROM python:3.11.8-alpine3.18

# 安装antiword及编译依赖
RUN apk add --no-cache antiword python3-dev gcc musl-dev

# 安装Python依赖
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt

# 复制应用代码
COPY . /app
WORKDIR /app

CMD ["python", "your_main_script.py"]

requirements.txt需包含以下依赖:

langchain
python-docx
textract

内容的提问来源于stack exchange,提问作者Eric Vaillancourt

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.24 10:23:18