如何使用MuPDF.js提取PDF高亮批注对应的文本?
使用MuPDF.js提取PDF高亮批注关联文本
我用MuPDF.js提取PDF批注时,能检测到高亮批注类型,但annot.hasRect()返回false,annot.getContents()也拿不到高亮对应的文本,原代码如下:
import * as fs from "fs"; import * as mupdfjs from "mupdf/mupdfjs"; async function extractAnnotations() { try { // Read the PDF file as a buffer let fileData = fs.readFileSync("Example.pdf"); // Open the document using mupdfjs.PDFDocument.openDocument let document = await mupdfjs.PDFDocument.openDocument( fileData, "application/pdf" ); // Loop through all pages of the document let pageCount = document.countPages(); // Get total number of pages let i = 0; while (i < pageCount) { const page = new mupdfjs.PDFPage(document, i); const annots = page.getAnnotations(); annots.forEach((annot, index) => { console.log(`Annotation ${index + 1}:`); console.log(` Has Rect: ${annot.hasRect()}`); console.log(` Contents: ${annot.getContents()}`); }); i++; } } catch (err) { console.error("Error extracting annotations:", err); } } extractAnnotations();
解决方法
高亮批注的文本提取需要注意两个关键点:
- 高亮批注的区域不是单一矩形,而是用四边形集合存储,需要用
annot.getQuadPoints()获取,而非getRect() - 需要调用页面的
getTextInQuad()方法,传入每个四边形区域来提取对应文本
修改后的完整代码:
import * as fs from "fs"; import * as mupdfjs from "mupdf/mupdfjs"; async function extractAnnotations() { try { let fileData = fs.readFileSync("Example.pdf"); let document = await mupdfjs.PDFDocument.openDocument( fileData, "application/pdf" ); let pageCount = document.countPages(); let i = 0; while (i < pageCount) { const page = new mupdfjs.PDFPage(document, i); const annots = page.getAnnotations(); annots.forEach((annot, index) => { // 仅处理高亮类型批注 if (annot.getType() === mupdfjs.PDFAnnotation.Type.Highlight) { console.log(`Highlight Annotation ${index + 1}:`); // 获取高亮的所有四边形区域 const quadPoints = annot.getQuadPoints(); if (quadPoints.length > 0) { let highlightedText = ""; // 遍历每个四边形,提取对应文本 quadPoints.forEach(quad => { const text = page.getTextInQuad(quad); highlightedText += text + " "; }); console.log(` Highlighted Text: ${highlightedText.trim()}`); } else { console.log(" No quad points found for this highlight."); } } else { // 其他类型批注的原有处理逻辑 console.log(`Annotation ${index + 1} (Type: ${annot.getType()}):`); console.log(` Has Rect: ${annot.hasRect()}`); console.log(` Contents: ${annot.getContents()}`); } }); i++; } } catch (err) { console.error("Error extracting annotations:", err); } } extractAnnotations();
关键说明
annot.getType()用于判断批注类型,mupdfjs.PDFAnnotation.Type.Highlight是MuPDF.js定义的高亮批注常量getQuadPoints()返回的是PDFQuad对象数组,每个对象代表高亮区域的一个四边形page.getTextInQuad(quad)会提取该四边形覆盖的页面文本,将多个四边形的文本拼接即可得到完整的高亮内容
内容的提问来源于stack exchange,提问作者Joe
相关产品推荐
相关产品推荐

