You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将自研爬虫获取的内容加载至LangChain的VectorstoreIndexCreator?

用户问题

我已经实现了一个可爬取指定URL及其子页面内容的函数,现希望将获取到的文本内容加载到LangChain的VectorstoreIndexCreator中。我在langchain.document_loaders中未找到合适的加载器,是否应使用BaseLoader?具体该如何操作?

原始代码

import requests
from bs4 import BeautifulSoup

import openai
from langchain.document_loaders.base import Document
from langchain.indexes import VectorstoreIndexCreator


def get_company_info_from_web(company_url: str, max_crawl_pages: int = 10, questions=None):

    # 访问URL并获取链接
    links = get_links_from_page(company_url)

    # get_text_content_from_page函数访问URL并返回文本与URL的元组
    for text, url in get_text_content_from_page(links[:max_crawl_pages]): 
        # 将文本内容(字符串)添加至索引
        # loader????

    index= VectorstoreIndexCreator().from_documents([Document(page_content=content, metadata={"source": url})])

    # 最后,向向量数据库发起查询:
    DEFAULT_QUERY = f"What does the company do? Who are key people in this company? Can you tell me contact information?"
    query = questions or DEFAULT_QUERY
    logger.info(f"Query: {query}")
    result = index.query_with_sources(query)

    logger.info(f"Result:\n {result['answer']}")
    logger.info(f"Sources:\n {result['sources']}")

    return result['answer'], result['sources']

解答

不需要使用BaseLoader,直接用Document类封装即可

你已经导入的Document类完全满足需求——VectorstoreIndexCreator.from_documents()方法可以直接接收Document对象的列表,没必要额外实现BaseLoader。BaseLoader是用于从自定义数据源(如特定API、小众本地文件格式)实现加载逻辑的场景,而你已经有了爬取好的文本内容,直接封装成Document是最高效的方式。

具体修改步骤

  1. 初始化空列表存储Document对象:在循环前创建列表,用来收集所有爬取页面的内容。
  2. 循环封装Document并添加到列表:把每个爬取到的文本和对应URL封装成Document,追加到列表中。
  3. 用完整的Document列表创建索引:将收集好的列表传入from_documents方法,避免只使用最后一个页面的内容。

修改后的完整代码

import requests
from bs4 import BeautifulSoup
import logging

import openai
from langchain.document_loaders.base import Document
from langchain.indexes import VectorstoreIndexCreator

# 初始化日志(原代码缺失此步骤)
logger = logging.getLogger(__name__)
logging.basicConfig(level=logging.INFO)

def get_company_info_from_web(company_url: str, max_crawl_pages: int = 10, questions=None):
    # 访问URL并获取链接
    links = get_links_from_page(company_url)

    # 初始化列表存储所有Document对象
    documents = []

    # 遍历爬取到的页面,封装成Document
    for text, url in get_text_content_from_page(links[:max_crawl_pages]): 
        doc = Document(page_content=text, metadata={"source": url})
        documents.append(doc)

    # 用所有Document创建索引
    index = VectorstoreIndexCreator().from_documents(documents)

    # 查询逻辑
    DEFAULT_QUERY = "What does the company do? Who are key people in this company? Can you tell me contact information?"
    query = questions or DEFAULT_QUERY
    logger.info(f"Query: {query}")
    result = index.query_with_sources(query)

    logger.info(f"Result:\n {result['answer']}")
    logger.info(f"Sources:\n {result['sources']}")

    return result['answer'], result['sources']

关键改动说明

  • 新增documents列表收集所有爬取页面的Document对象,修复原代码只使用最后一个页面内容的问题。
  • 修正原代码中content变量未定义的错误,改用循环中的text变量。
  • 补充日志初始化代码,避免原代码直接调用logger报错。

内容的提问来源于stack exchange,提问作者PetrSevcik

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 21:48:20