如何在ReactJS中根据行号高亮PDF页面文本
在ReactJS中根据行号高亮PDF文本的实现方案
要实现基于行号的PDF文本高亮,我们可以借助@react-pdf-viewer/core和@react-pdf-viewer/highlight插件,结合PDF.js的文本提取能力完成。以下是具体实现步骤:
1. 安装依赖
首先安装高亮插件并引入样式:
npm install @react-pdf-viewer/highlight
import "@react-pdf-viewer/highlight/lib/styles/index.css";
2. 完整实现代码
替换你的App组件为以下代码,核心逻辑是在PDF加载完成后提取页面文本,匹配目标行号对应的文本并生成高亮:
import { Worker, Viewer, useDocument } from "@react-pdf-viewer/core"; import { defaultLayoutPlugin } from "@react-pdf-viewer/default-layout"; import { highlightPlugin, HighlightArea, HighlightTarget } from "@react-pdf-viewer/highlight"; import "@react-pdf-viewer/core/lib/styles/index.css"; import "@react-pdf-viewer/default-layout/lib/styles/index.css"; import "@react-pdf-viewer/highlight/lib/styles/index.css"; import { useEffect, useState } from "react"; export default function App() { const data = { text: "The requirements include...", sourceDocuments: [ { pageContent:"Functionality requirements, backend functionality\nUser data information\nAbility to collect and sort...", metadata: { "loc.lines.from": 161, "loc.lines.to": 173 } }, ], }; const defaultLayoutPluginInstance = defaultLayoutPlugin(); const [highlights, setHighlights] = useState<HighlightArea[]>([]); const { document } = useDocument(); useEffect(() => { const generateHighlights = async () => { if (!document) return; const newHighlights: HighlightArea[] = []; const targetDoc = data.sourceDocuments[0]; const startLine = targetDoc.metadata["loc.lines.from"]; const endLine = targetDoc.metadata["loc.lines.to"]; // 遍历PDF所有页面(可根据实际情况限定目标页面) for (let pageNum = 1; pageNum <= document.numPages; pageNum++) { const page = await document.getPage(pageNum); const textContent = await page.getTextContent(); // 拆分页面文本为行(基于文本块的y坐标变化判断换行) let currentLine = 1; let lineText = ""; const lines: { text: string; startIndex: number; endIndex: number }[] = []; textContent.items.forEach((item: any, index) => { const itemText = item.str; // 文本块y坐标下降代表换行 if (index > 0 && item.transform[5] < textContent.items[index - 1].transform[5]) { lines.push({ text: lineText.trim(), startIndex: index - lines.length, endIndex: index }); lineText = itemText; currentLine++; } else { lineText += " " + itemText; } }); // 添加最后一行 lines.push({ text: lineText.trim(), startIndex: textContent.items.length - lines.length, endIndex: textContent.items.length }); // 匹配目标行范围并生成高亮区域 if (currentLine >= startLine && currentLine <= endLine) { for (let lineIdx = startLine - 1; lineIdx <= endLine - 1; lineIdx++) { if (lineIdx >= lines.length) break; const targetLine = lines[lineIdx]; const startItem = textContent.items[targetLine.startIndex]; const endItem = textContent.items[targetLine.endIndex - 1]; newHighlights.push({ pageIndex: pageNum - 1, // PDF页面索引从0开始 rects: [ { x: startItem.transform[4], y: startItem.transform[5] - startItem.height, width: endItem.transform[4] + endItem.width - startItem.transform[4], height: startItem.height, }, ], content: targetLine.text, color: "#ffeb3b", // 自定义高亮颜色 }); } } } setHighlights(newHighlights); }; generateHighlights(); }, [document]); // 配置高亮插件 const highlightPluginInstance = highlightPlugin({ highlights: () => highlights, renderHighlightTarget: (props) => ( <HighlightTarget {...props}> {({ isHovered }) => ( <div style={{ backgroundColor: props.highlightArea.color, opacity: isHovered ? 0.6 : 0.3, cursor: "default", }} /> )} </HighlightTarget> ), }); return ( <Worker workerUrl="https://unpkg.com/pdfjs-dist@3.4.120/build/pdf.worker.min.js"> <div style={{ height: "720px" }}> <Viewer fileUrl="Brief.pdf" plugins={[defaultLayoutPluginInstance, highlightPluginInstance]} /> </div> </Worker> ); }
关键说明
- 行号匹配逻辑:通过PDF.js的
getTextContent()获取页面文本块,根据文本块的y坐标变化判断行分隔,拆分出每一行文本和对应的行号。 - 高亮区域生成:找到目标行对应的文本块后,计算其在页面上的坐标范围,生成
HighlightArea对象,通过高亮插件渲染高亮层。 - 适配调整:如果PDF的行号是基于原始文档的排版行,可能需要用
pageContent和PDF提取的文本做精准匹配,避免因PDF排版差异导致行号错位。
内容的提问来源于stack exchange,提问作者Venkatesh
相关产品推荐
相关产品推荐

