You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在LangChain的Confluence Loader中自定义expand参数?

优化LangChain Confluence Loader的expand参数传递方案

问题背景

使用LangChain的Confluence Loader时,需要自定义Confluence API的expand参数来获取额外字段,但load()方法未提供直接传入expand的入口,而是通过ContentFormat枚举的value传递该参数。目前通过重写ContentFormat枚举修改STORAGE值能实现需求,但不够优雅,以下是几种更合理的解决方案:


方案1:扩展原ContentFormat枚举(推荐)

无需完全重写枚举,直接在原ContentFormat上添加自定义项,保留原有逻辑的同时实现扩展:

from langchain.document_loaders.confluence import ContentFormat

# 添加自定义扩展项,指定需要的expand字段
ContentFormat.CUSTOM_STORAGE = "body.storage,version"

# 为自定义项兼容get_content方法(确保能正确提取内容)
def custom_get_content(self, page: dict) -> str:
    return page["body"]["storage"]["value"]

# 绑定方法到枚举实例
ContentFormat.CUSTOM_STORAGE.get_content = custom_get_content.__get__(ContentFormat.CUSTOM_STORAGE, ContentFormat)

调用时直接指定该自定义枚举值:

documents = loader.load(
    space_key="KB",
    include_attachments=False,
    keep_newlines=True,
    keep_markdown_format=True,
    content_format=ContentFormat.CUSTOM_STORAGE
)

方案2:子类化ConfluenceLoader添加expand参数

通过扩展ConfluenceLoader类,给load()方法新增expand参数,让调用更直观:

from langchain.document_loaders.confluence import ConfluenceLoader, ContentFormat
from langchain.schema import Document
from typing import Optional, List

class CustomConfluenceLoader(ConfluenceLoader):
    def load(
        self,
        space_key: Optional[str] = None,
        page_ids: Optional[List[str]] = None,
        label: Optional[str] = None,
        cql: Optional[str] = None,
        include_restricted_content: bool = False,
        include_archived_content: bool = False,
        include_attachments: bool = False,
        include_comments: bool = False,
        content_format: ContentFormat = ContentFormat.STORAGE,
        limit: Optional[int] = 50,
        max_pages: Optional[int] = 1000,
        ocr_languages: Optional[str] = None,
        keep_markdown_format: bool = False,
        keep_newlines: bool = False,
        expand: Optional[str] = None,  # 新增自定义expand参数
    ) -> List[Document]:
        docs = []
        # 优先使用自定义expand,否则沿用content_format的value
        target_expand = expand if expand is not None else content_format.value
        
        if space_key:
            pages = self.paginate_request(
                self.confluence.get_all_pages_from_space,
                space=space_key,
                limit=limit,
                max_pages=max_pages,
                status="any" if include_archived_content else "current",
                expand=target_expand,
            )
            docs += self.process_pages(
                pages,
                include_restricted_content,
                include_attachments,
                include_comments,
                content_format,
                ocr_languages=ocr_languages,
                keep_markdown_format=keep_markdown_format,
                keep_newlines=keep_newlines,
            )
        # 复制原load方法中处理page_ids、label、cql的逻辑
        if page_ids:
            pages = [self.confluence.get_page_by_id(page_id, expand=target_expand) for page_id in page_ids]
            docs += self.process_pages(
                pages,
                include_restricted_content,
                include_attachments,
                include_comments,
                content_format,
                ocr_languages=ocr_languages,
                keep_markdown_format=keep_markdown_format,
                keep_newlines=keep_newlines,
            )
        if label:
            pages = self.paginate_request(
                self.confluence.get_all_pages_by_label,
                label=label,
                limit=limit,
                max_pages=max_pages,
                expand=target_expand,
            )
            docs += self.process_pages(
                pages,
                include_restricted_content,
                include_attachments,
                include_comments,
                content_format,
                ocr_languages=ocr_languages,
                keep_markdown_format=keep_markdown_format,
                keep_newlines=keep_newlines,
            )
        if cql:
            pages = self.paginate_request(
                self.confluence.cql,
                cql=cql,
                limit=limit,
                max_pages=max_pages,
                expand=target_expand,
            )
            docs += self.process_pages(
                pages,
                include_restricted_content,
                include_attachments,
                include_comments,
                content_format,
                ocr_languages=ocr_languages,
                keep_markdown_format=keep_markdown_format,
                keep_newlines=keep_newlines,
            )
        return docs

使用实例:

# 初始化自定义Loader
loader = CustomConfluenceLoader(url="你的Confluence地址", username="用户名", api_key="API密钥")
# 调用时直接传入expand参数
documents = loader.load(
    space_key="KB",
    include_attachments=False,
    keep_newlines=True,
    keep_markdown_format=True,
    expand="body.storage,version"
)

方案3:绕过load方法,直接调用底层API

如果只是临时需求,可直接调用Loader的底层方法,完全自主控制参数:

# 获取原始页面数据,直接指定expand
pages = loader.confluence.get_all_pages_from_space(
    space="KB",
    limit=50,
    max_pages=1000,
    status="current",
    expand="body.storage,version"
)
# 处理成LangChain Document对象
documents = loader.process_pages(
    pages,
    include_restricted_content=False,
    include_attachments=False,
    include_comments=False,
    content_format=ContentFormat.STORAGE,
    keep_markdown_format=True,
    keep_newlines=True
)

内容的提问来源于stack exchange,提问作者Tim Schill

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.07 19:50:34