You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何通过pdf-lib获取PDF表单字段所在页码(React TypeScript场景)

在React TypeScript中用pdf-lib获取表单字段及对应页码的解决方案

针对你提出的两个需求,以下是基于pdf-lib的具体实现方案:

1. 从PDFField对象获取对应PDFPage页码

核心逻辑是先建立字段引用与页码的映射关系——遍历所有页面收集其注解(Annots),再通过字段引用匹配对应的页面。

2. 获取PDFDocument中所有文本字段及对应页码

结合上述映射,筛选出所有文本类型的表单字段,并关联其所在页码。

完整实现代码

import { PDFDocumentFactory, PDFName, PDFRef } from 'pdf-lib';
import fs from 'fs';

// 获取AcroForm中的所有字段引用
const getAcroFieldRefs = (pdfDoc: any): PDFRef[] => {
  const acroFormRef = pdfDoc.catalog.getMaybe('AcroForm');
  if (!acroFormRef) return [];
  
  const acroForm = pdfDoc.index.lookup(acroFormRef);
  const fieldsRef = acroForm.getMaybe('Fields');
  if (!fieldsRef) return [];
  
  const acroFields = pdfDoc.index.lookup(fieldsRef);
  return acroFields.array as PDFRef[];
};

// 构建字段引用到页码的映射
const buildFieldToPageMap = (pdfDoc: any): Map<string, number> => {
  const fieldPageMap = new Map<string, number>();
  const pages = pdfDoc.getPages();

  pages.forEach((page: any, pageIndex: number) => {
    const annotsRaw = pdfDoc.index.lookupMaybe(page.getMaybe('Annots'));
    if (!annotsRaw) return;
    
    const annots = annotsRaw.array as PDFRef[];
    annots.forEach(annotRef => {
      fieldPageMap.set(annotRef.toString(), pageIndex + 1); // 页码从1开始计数
    });
  });

  return fieldPageMap;
};

// 获取所有文本字段及对应页码
const getTextFieldsWithPageNumbers = (pdfDoc: any) => {
  const fieldRefs = getAcroFieldRefs(pdfDoc);
  const fieldPageMap = buildFieldToPageMap(pdfDoc);
  
  return fieldRefs.map(fieldRef => {
    const field = pdfDoc.index.lookup(fieldRef);
    // 判断是否为文本字段(FT属性为Tx)
    const fieldType = field.getMaybe('FT');
    if (fieldType instanceof PDFName && fieldType.value === 'Tx') {
      // 获取字段名称
      const fieldName = field.getMaybe('T')?.value || '未命名字段';
      // 获取对应页码
      const pageNumber = fieldPageMap.get(fieldRef.toString()) || -1;
      
      return {
        fieldName,
        fieldRef: fieldRef.toString(),
        pageNumber,
        field // 原始PDFField对象
      };
    }
    return null;
  }).filter(Boolean); // 过滤非文本字段
};

// 示例使用
const pdfDoc = PDFDocumentFactory.load(fs.readFileSync('./form.pdf'));
const textFields = getTextFieldsWithPageNumbers(pdfDoc);

console.log('所有文本字段及页码:', textFields);

// 单独获取某个PDFField的页码示例
const sampleFieldRef = getAcroFieldRefs(pdfDoc)[0];
const samplePageNumber = buildFieldToPageMap(pdfDoc).get(sampleFieldRef.toString());
console.log(`字段${sampleFieldRef.toString()}所在页码: ${samplePageNumber}`);

关键逻辑说明

  • 字段-页码映射构建:遍历所有页面,将每个页面的注解(表单字段本质是页面注解)引用与页码关联,存入Map中。
  • 文本字段筛选:通过字段的FT属性判断类型,Tx代表文本字段。
  • 页码查询:通过字段的引用字符串,直接从Map中取出对应的页码。

内容的提问来源于stack exchange,提问作者Shahid Nauman

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.25 05:40:41