You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用LlamaIndex与LangChain同时索引PDF文本和表格(含OpenAI密钥)

解决PDF文本与表格同时索引的方案

核心思路

替换原仅提取纯文本的PDFReader,改用支持表格结构化解析的加载器,将表格内容转为大模型易识别的格式(如Markdown表格)后,再传入LlamaIndex构建索引,让问答机器人能识别并检索表格数据。

具体实现方案

方案一:使用LlamaHub的PyMuPDFLoader

PyMuPDF(基于fitz库)能精准提取PDF中的表格结构,并自动转为Markdown格式,适配LlamaIndex的索引逻辑。代码替换如下:

from llama_index import download_loader
from pathlib import Path
import os
from llama_index import GPTSimpleVectorIndex, LLMPredictor, ServiceContext
from langchain.llms import OpenAI
import logging

logger = logging.getLogger(__name__)
INDEX_FILE = "index.json"

def ask(file):
    print(" Loading...")
    # 替换为支持表格解析的PyMuPDFLoader
    PyMuPDFLoader = download_loader("PyMuPDFLoader")
    loader = PyMuPDFLoader()
    documents = loader.load_data(file=Path(file))
    print("Path: ", Path(file))

    # 原有索引加载/生成逻辑保持不变
    if os.path.exists(INDEX_FILE):
        logger.info("found index.json in the directory")
        index = GPTSimpleVectorIndex.load_from_disk(INDEX_FILE)
    else:
        logger.info("didnt find index.json in the directory")
        llm_predictor = LLMPredictor(llm=OpenAI(temperature=0, model_name="text-davinci-003"))
        service_context = ServiceContext.from_defaults(llm_predictor=llm_predictor, chunk_size_limit=1024)
        index = GPTSimpleVectorIndex.from_documents(documents, service_context=service_context)
        index.save_to_disk(INDEX_FILE)

方案二:结合LangChain的UnstructuredFileLoader

通过LangChain的UnstructuredFileLoader,以元素拆分模式(mode="elements")将PDF拆分为文本、表格等独立元素,再转为LlamaIndex兼容的文档格式:

from langchain.document_loaders import UnstructuredFileLoader
from llama_index import LangchainReader, GPTSimpleVectorIndex, LLMPredictor, ServiceContext
from langchain.llms import OpenAI
import os
import logging

logger = logging.getLogger(__name__)
INDEX_FILE = "index.json"

def ask(file):
    print(" Loading...")
    # 启用元素模式加载PDF,单独识别表格
    loader = UnstructuredFileLoader(file, mode="elements")
    langchain_docs = loader.load()
    # 将LangChain文档转换为LlamaIndex可处理的格式
    llama_reader = LangchainReader()
    documents = llama_reader.load_langchain_documents(langchain_docs)
    print("Path: ", Path(file))

    # 原有索引加载/生成逻辑保持不变
    if os.path.exists(INDEX_FILE):
        logger.info("found index.json in the directory")
        index = GPTSimpleVectorIndex.load_from_disk(INDEX_FILE)
    else:
        logger.info("didnt find index.json in the directory")
        llm_predictor = LLMPredictor(llm=OpenAI(temperature=0, model_name="text-davinci-003"))
        service_context = ServiceContext.from_defaults(llm_predictor=llm_predictor, chunk_size_limit=1024)
        index = GPTSimpleVectorIndex.from_documents(documents, service_context=service_context)
        index.save_to_disk(INDEX_FILE)

优化索引效果的细节

  • 保持表格的Markdown格式:两种方案都会自动将表格转为Markdown,这种结构化格式能让大模型更好理解表格的行列关系。
  • 调整分块大小:如果表格内容较长,可将chunk_size_limit调至2048,避免表格被拆分得过于零散,破坏数据完整性。
  • 针对性检索(可选):可自定义节点处理逻辑,给表格类型的文档节点添加标签,在检索时优先匹配表格类内容,提升查询精度。

内容的提问来源于stack exchange,提问作者Harshit

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.22 21:02:32