如何用LlamaIndex与LangChain同时索引PDF文本和表格(含OpenAI密钥)
解决PDF文本与表格同时索引的方案
核心思路
替换原仅提取纯文本的PDFReader,改用支持表格结构化解析的加载器,将表格内容转为大模型易识别的格式(如Markdown表格)后,再传入LlamaIndex构建索引,让问答机器人能识别并检索表格数据。
具体实现方案
方案一:使用LlamaHub的PyMuPDFLoader
PyMuPDF(基于fitz库)能精准提取PDF中的表格结构,并自动转为Markdown格式,适配LlamaIndex的索引逻辑。代码替换如下:
from llama_index import download_loader from pathlib import Path import os from llama_index import GPTSimpleVectorIndex, LLMPredictor, ServiceContext from langchain.llms import OpenAI import logging logger = logging.getLogger(__name__) INDEX_FILE = "index.json" def ask(file): print(" Loading...") # 替换为支持表格解析的PyMuPDFLoader PyMuPDFLoader = download_loader("PyMuPDFLoader") loader = PyMuPDFLoader() documents = loader.load_data(file=Path(file)) print("Path: ", Path(file)) # 原有索引加载/生成逻辑保持不变 if os.path.exists(INDEX_FILE): logger.info("found index.json in the directory") index = GPTSimpleVectorIndex.load_from_disk(INDEX_FILE) else: logger.info("didnt find index.json in the directory") llm_predictor = LLMPredictor(llm=OpenAI(temperature=0, model_name="text-davinci-003")) service_context = ServiceContext.from_defaults(llm_predictor=llm_predictor, chunk_size_limit=1024) index = GPTSimpleVectorIndex.from_documents(documents, service_context=service_context) index.save_to_disk(INDEX_FILE)
方案二:结合LangChain的UnstructuredFileLoader
通过LangChain的UnstructuredFileLoader,以元素拆分模式(mode="elements")将PDF拆分为文本、表格等独立元素,再转为LlamaIndex兼容的文档格式:
from langchain.document_loaders import UnstructuredFileLoader from llama_index import LangchainReader, GPTSimpleVectorIndex, LLMPredictor, ServiceContext from langchain.llms import OpenAI import os import logging logger = logging.getLogger(__name__) INDEX_FILE = "index.json" def ask(file): print(" Loading...") # 启用元素模式加载PDF,单独识别表格 loader = UnstructuredFileLoader(file, mode="elements") langchain_docs = loader.load() # 将LangChain文档转换为LlamaIndex可处理的格式 llama_reader = LangchainReader() documents = llama_reader.load_langchain_documents(langchain_docs) print("Path: ", Path(file)) # 原有索引加载/生成逻辑保持不变 if os.path.exists(INDEX_FILE): logger.info("found index.json in the directory") index = GPTSimpleVectorIndex.load_from_disk(INDEX_FILE) else: logger.info("didnt find index.json in the directory") llm_predictor = LLMPredictor(llm=OpenAI(temperature=0, model_name="text-davinci-003")) service_context = ServiceContext.from_defaults(llm_predictor=llm_predictor, chunk_size_limit=1024) index = GPTSimpleVectorIndex.from_documents(documents, service_context=service_context) index.save_to_disk(INDEX_FILE)
优化索引效果的细节
- 保持表格的Markdown格式:两种方案都会自动将表格转为Markdown,这种结构化格式能让大模型更好理解表格的行列关系。
- 调整分块大小:如果表格内容较长,可将
chunk_size_limit调至2048,避免表格被拆分得过于零散,破坏数据完整性。 - 针对性检索(可选):可自定义节点处理逻辑,给表格类型的文档节点添加标签,在检索时优先匹配表格类内容,提升查询精度。
内容的提问来源于stack exchange,提问作者Harshit
相关产品推荐
相关产品推荐

