You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用pdfjsLib提取PDF表格的水平图形线条?

解决方案:用pdfjsLib提取PDF表格的水平图形线条

getAnnotations()返回空是因为表格的图形线条不属于注解(Annotations),而是页面内容流中的路径绘制元素(比如直线、矩形)。要提取这些线条,需要解析页面的操作符列表(Operator List),从中筛选水平线段。

实现步骤与代码示例

async function extractHorizontalLines(pdfUrl) {
  const loadingTask = pdfjsLib.getDocument(pdfUrl);
  const pdf = await loadingTask.promise;
  const horizontalLines = [];

  for (let pageNum = 1; pageNum <= pdf.numPages; pageNum++) {
    const page = await pdf.getPage(pageNum);
    const operatorList = await page.getOperatorList();
    const viewport = page.getViewport({ scale: 1 });
    const pathData = [];
    let currentPath = [];

    // 遍历页面操作符,收集所有路径数据
    for (let i = 0; i < operatorList.fnArray.length; i++) {
      const operator = operatorList.fnArray[i];
      const args = operatorList.argsArray[i];

      switch (operator) {
        case pdfjsLib.OPS.moveTo: // 路径起点
          currentPath.push({ x: args[0], y: args[1] });
          break;
        case pdfjsLib.OPS.lineTo: // 线段到下一点
          currentPath.push({ x: args[0], y: args[1] });
          break;
        case pdfjsLib.OPS.closePath: // 闭合路径
          currentPath.push(currentPath[0]);
          pathData.push([...currentPath]);
          currentPath = [];
          break;
        case pdfjsLib.OPS.stroke: // 描边路径(确认路径)
        case pdfjsLib.OPS.strokeAndFill:
        case pdfjsLib.OPS.fillAndStroke:
          if (currentPath.length > 0) {
            pathData.push([...currentPath]);
            currentPath = [];
          }
          break;
        case pdfjsLib.OPS.rectangle: // 矩形(表格边框常用此操作)
          const [x, y, width, height] = args;
          // 拆解矩形的四条边
          const rectPath = [
            { x, y },
            { x: x + width, y },
            { x: x + width, y: y + height },
            { x, y: y + height },
            { x, y }
          ];
          pathData.push(rectPath);
          break;
      }
    }

    // 从路径中筛选水平线条
    for (const path of pathData) {
      if (path.length < 2) continue;
      for (let j = 0; j < path.length - 1; j++) {
        const p1 = path[j];
        const p2 = path[j + 1];
        // 水平线条判断:y坐标误差小于1e-3,且x坐标有差异(排除单点)
        if (Math.abs(p1.y - p2.y) < 1e-3 && Math.abs(p1.x - p2.x) > 1e-3) {
          // 转换为页面视口坐标(可选,根据需求调整)
          const vp1 = viewport.convertToViewportPoint(p1.x, p1.y);
          const vp2 = viewport.convertToViewportPoint(p2.x, p2.y);
          horizontalLines.push({
            page: pageNum,
            xStart: Math.min(vp1.x, vp2.x),
            xEnd: Math.max(vp1.x, vp2.x),
            y: vp1.y,
            originalCoords: { p1, p2 }
          });
        }
      }
    }
  }

  return horizontalLines;
}

// 调用示例
extractHorizontalLines('your-document.pdf').then(lines => {
  console.log('提取到的水平线条:', lines);
});

关键说明

  1. 操作符解析:页面的所有绘制逻辑都在operatorList中,我们关注路径相关的操作(移动、画线、矩形、描边等),这些是图形线条的来源。
  2. 矩形处理:很多PDF表格边框是用矩形绘制的,所以单独处理rectangle操作,拆解其四条边来筛选水平线段。
  3. 浮点精度:PDF坐标是浮点数,比较时用1e-3的误差范围避免精度问题。
  4. 视口转换:convertToViewportPoint会把PDF内部坐标转换为页面显示坐标,适合前端渲染场景;如果需要原始PDF坐标,可以跳过这一步。

额外注意事项

  • 如果表格线条是用极窄的填充矩形模拟的(比如高度<1的矩形),可以在筛选时增加判断:矩形高度小于阈值,且宽度较大,将其视为水平线条。
  • 页面旋转会被viewport自动处理,转换后的坐标是正确的显示坐标。

内容的提问来源于stack exchange,提问作者jsdbt

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.27 02:12:08