如何在LangChain的Confluence Loader中自定义expand参数?
优化LangChain Confluence Loader的expand参数传递方案
问题背景
使用LangChain的Confluence Loader时,需要自定义Confluence API的expand参数来获取额外字段,但load()方法未提供直接传入expand的入口,而是通过ContentFormat枚举的value传递该参数。目前通过重写ContentFormat枚举修改STORAGE值能实现需求,但不够优雅,以下是几种更合理的解决方案:
方案1:扩展原ContentFormat枚举(推荐)
无需完全重写枚举,直接在原ContentFormat上添加自定义项,保留原有逻辑的同时实现扩展:
from langchain.document_loaders.confluence import ContentFormat # 添加自定义扩展项,指定需要的expand字段 ContentFormat.CUSTOM_STORAGE = "body.storage,version" # 为自定义项兼容get_content方法(确保能正确提取内容) def custom_get_content(self, page: dict) -> str: return page["body"]["storage"]["value"] # 绑定方法到枚举实例 ContentFormat.CUSTOM_STORAGE.get_content = custom_get_content.__get__(ContentFormat.CUSTOM_STORAGE, ContentFormat)
调用时直接指定该自定义枚举值:
documents = loader.load( space_key="KB", include_attachments=False, keep_newlines=True, keep_markdown_format=True, content_format=ContentFormat.CUSTOM_STORAGE )
方案2:子类化ConfluenceLoader添加expand参数
通过扩展ConfluenceLoader类,给load()方法新增expand参数,让调用更直观:
from langchain.document_loaders.confluence import ConfluenceLoader, ContentFormat from langchain.schema import Document from typing import Optional, List class CustomConfluenceLoader(ConfluenceLoader): def load( self, space_key: Optional[str] = None, page_ids: Optional[List[str]] = None, label: Optional[str] = None, cql: Optional[str] = None, include_restricted_content: bool = False, include_archived_content: bool = False, include_attachments: bool = False, include_comments: bool = False, content_format: ContentFormat = ContentFormat.STORAGE, limit: Optional[int] = 50, max_pages: Optional[int] = 1000, ocr_languages: Optional[str] = None, keep_markdown_format: bool = False, keep_newlines: bool = False, expand: Optional[str] = None, # 新增自定义expand参数 ) -> List[Document]: docs = [] # 优先使用自定义expand,否则沿用content_format的value target_expand = expand if expand is not None else content_format.value if space_key: pages = self.paginate_request( self.confluence.get_all_pages_from_space, space=space_key, limit=limit, max_pages=max_pages, status="any" if include_archived_content else "current", expand=target_expand, ) docs += self.process_pages( pages, include_restricted_content, include_attachments, include_comments, content_format, ocr_languages=ocr_languages, keep_markdown_format=keep_markdown_format, keep_newlines=keep_newlines, ) # 复制原load方法中处理page_ids、label、cql的逻辑 if page_ids: pages = [self.confluence.get_page_by_id(page_id, expand=target_expand) for page_id in page_ids] docs += self.process_pages( pages, include_restricted_content, include_attachments, include_comments, content_format, ocr_languages=ocr_languages, keep_markdown_format=keep_markdown_format, keep_newlines=keep_newlines, ) if label: pages = self.paginate_request( self.confluence.get_all_pages_by_label, label=label, limit=limit, max_pages=max_pages, expand=target_expand, ) docs += self.process_pages( pages, include_restricted_content, include_attachments, include_comments, content_format, ocr_languages=ocr_languages, keep_markdown_format=keep_markdown_format, keep_newlines=keep_newlines, ) if cql: pages = self.paginate_request( self.confluence.cql, cql=cql, limit=limit, max_pages=max_pages, expand=target_expand, ) docs += self.process_pages( pages, include_restricted_content, include_attachments, include_comments, content_format, ocr_languages=ocr_languages, keep_markdown_format=keep_markdown_format, keep_newlines=keep_newlines, ) return docs
使用实例:
# 初始化自定义Loader loader = CustomConfluenceLoader(url="你的Confluence地址", username="用户名", api_key="API密钥") # 调用时直接传入expand参数 documents = loader.load( space_key="KB", include_attachments=False, keep_newlines=True, keep_markdown_format=True, expand="body.storage,version" )
方案3:绕过load方法,直接调用底层API
如果只是临时需求,可直接调用Loader的底层方法,完全自主控制参数:
# 获取原始页面数据,直接指定expand pages = loader.confluence.get_all_pages_from_space( space="KB", limit=50, max_pages=1000, status="current", expand="body.storage,version" ) # 处理成LangChain Document对象 documents = loader.process_pages( pages, include_restricted_content=False, include_attachments=False, include_comments=False, content_format=ContentFormat.STORAGE, keep_markdown_format=True, keep_newlines=True )
内容的提问来源于stack exchange,提问作者Tim Schill
相关产品推荐
相关产品推荐

