You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Langchain-JS中UnstructuredDirectoryLoader加载大PDF超时问题求助

解决方案:处理大型PDF时Unstructured API hi_res策略超时问题

1. 拆分PDF为小批次处理

600页的PDF单次处理负载过高,优先拆分后分批处理再合并结果。可使用pdf-lib库在NextJS中拆分PDF:

import { PDFDocument } from 'pdf-lib';

async function splitPdfIntoChunks(pdfBuffer: Buffer, chunkSize: number = 50) {
  const pdfDoc = await PDFDocument.load(pdfBuffer);
  const totalPages = pdfDoc.getPageCount();
  const chunks: Buffer[] = [];

  for (let i = 0; i < totalPages; i += chunkSize) {
    const end = Math.min(i + chunkSize, totalPages);
    const newPdfDoc = await PDFDocument.create();
    
    // 复制对应页码的页面
    const pages = await newPdfDoc.copyPages(pdfDoc, Array.from({length: end - i}, (_, idx) => i + idx));
    pages.forEach(page => newPdfDoc.addPage(page));
    
    const chunkBuffer = await newPdfDoc.save();
    chunks.push(chunkBuffer);
  }
  
  return chunks;
}

之后对每个拆分后的PDF块单独使用hi_res策略加载,最后合并所有提取结果:

import { UnstructuredLoader } from "langchain/document_loaders/fs/unstructured";

async function processPdfChunks(chunks: Buffer[]) {
  const allDocuments = [];
  
  for (const chunk of chunks) {
    const loader = new UnstructuredLoader(chunk, {
      apiUrl: "YOUR_UNSTRUCTURED_API_URL",
      strategy: "hi_res",
      timeout: 300000, // 5分钟超时
      elementsToExtract: ["Text", "Table"], // 只提取需要的元素,减少处理量
    });
    
    const docs = await loader.load();
    allDocuments.push(...docs);
  }
  
  return allDocuments;
}

2. 调整NextJS服务端超时限制

如果用Vercel部署,Serverless Functions默认超时10秒(Pro计划90秒),远不足以处理大文件的hi_res解析,可通过以下方式解决:

  • 改用Edge Functions:Vercel Edge Functions超时可达60秒,更适配IO密集型任务;
  • 自建后端服务:用Express等框架部署独立服务,自定义超时设置(比如server.timeout = 300000),规避NextJS的Serverless超时限制;
  • 使用Vercel的Long-Running Functions(需Enterprise计划),支持最长1小时超时。

3. 本地部署Unstructured服务

云端Unstructured API通常有严格的超时和资源限制,本地部署Docker容器可完全自定义配置:

  1. 拉取Unstructured镜像:
docker pull downloads.unstructured.io/unstructured-io/unstructured:latest
  1. 启动容器,设置更长超时和充足资源:
docker run -p 8000:8000 --env TIMEOUT=300 --memory=16g downloads.unstructured.io/unstructured-io/unstructured:latest
  1. 在Langchain中使用本地API地址:
const loader = new UnstructuredLoader("path/to/large.pdf", {
  apiUrl: "http://localhost:8000/general/v0/general",
  strategy: "hi_res",
  timeout: 300000,
});

4. 优化Unstructured处理参数

通过减少不必要的处理步骤降低耗时:

  • 指定elementsToExtract:只提取你需要的元素(比如表格、文本),跳过图表等非必要内容;
  • 设置skipInferTableTypes:如果不需要识别特定类型的表格(如Excel嵌入表),可开启此参数减少计算;
  • 禁用OCR(如果PDF本身是可搜索的):
const loader = new UnstructuredLoader(pdfBuffer, {
  strategy: "hi_res",
  partitionConfig: {
    ocr_languages: [], // 禁用OCR
    skip_infer_table_types: true,
  },
});

内容的提问来源于stack exchange,提问作者graylagx2

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.10 20:40:12