如何通过pdf-lib获取PDF表单字段所在页码(React TypeScript场景)
在React TypeScript中用pdf-lib获取表单字段及对应页码的解决方案
针对你提出的两个需求,以下是基于pdf-lib的具体实现方案:
1. 从PDFField对象获取对应PDFPage页码
核心逻辑是先建立字段引用与页码的映射关系——遍历所有页面收集其注解(Annots),再通过字段引用匹配对应的页面。
2. 获取PDFDocument中所有文本字段及对应页码
结合上述映射,筛选出所有文本类型的表单字段,并关联其所在页码。
完整实现代码
import { PDFDocumentFactory, PDFName, PDFRef } from 'pdf-lib'; import fs from 'fs'; // 获取AcroForm中的所有字段引用 const getAcroFieldRefs = (pdfDoc: any): PDFRef[] => { const acroFormRef = pdfDoc.catalog.getMaybe('AcroForm'); if (!acroFormRef) return []; const acroForm = pdfDoc.index.lookup(acroFormRef); const fieldsRef = acroForm.getMaybe('Fields'); if (!fieldsRef) return []; const acroFields = pdfDoc.index.lookup(fieldsRef); return acroFields.array as PDFRef[]; }; // 构建字段引用到页码的映射 const buildFieldToPageMap = (pdfDoc: any): Map<string, number> => { const fieldPageMap = new Map<string, number>(); const pages = pdfDoc.getPages(); pages.forEach((page: any, pageIndex: number) => { const annotsRaw = pdfDoc.index.lookupMaybe(page.getMaybe('Annots')); if (!annotsRaw) return; const annots = annotsRaw.array as PDFRef[]; annots.forEach(annotRef => { fieldPageMap.set(annotRef.toString(), pageIndex + 1); // 页码从1开始计数 }); }); return fieldPageMap; }; // 获取所有文本字段及对应页码 const getTextFieldsWithPageNumbers = (pdfDoc: any) => { const fieldRefs = getAcroFieldRefs(pdfDoc); const fieldPageMap = buildFieldToPageMap(pdfDoc); return fieldRefs.map(fieldRef => { const field = pdfDoc.index.lookup(fieldRef); // 判断是否为文本字段(FT属性为Tx) const fieldType = field.getMaybe('FT'); if (fieldType instanceof PDFName && fieldType.value === 'Tx') { // 获取字段名称 const fieldName = field.getMaybe('T')?.value || '未命名字段'; // 获取对应页码 const pageNumber = fieldPageMap.get(fieldRef.toString()) || -1; return { fieldName, fieldRef: fieldRef.toString(), pageNumber, field // 原始PDFField对象 }; } return null; }).filter(Boolean); // 过滤非文本字段 }; // 示例使用 const pdfDoc = PDFDocumentFactory.load(fs.readFileSync('./form.pdf')); const textFields = getTextFieldsWithPageNumbers(pdfDoc); console.log('所有文本字段及页码:', textFields); // 单独获取某个PDFField的页码示例 const sampleFieldRef = getAcroFieldRefs(pdfDoc)[0]; const samplePageNumber = buildFieldToPageMap(pdfDoc).get(sampleFieldRef.toString()); console.log(`字段${sampleFieldRef.toString()}所在页码: ${samplePageNumber}`);
关键逻辑说明
- 字段-页码映射构建:遍历所有页面,将每个页面的注解(表单字段本质是页面注解)引用与页码关联,存入Map中。
- 文本字段筛选:通过字段的
FT属性判断类型,Tx代表文本字段。 - 页码查询:通过字段的引用字符串,直接从Map中取出对应的页码。
内容的提问来源于stack exchange,提问作者Shahid Nauman
相关产品推荐
相关产品推荐

