如何借助Azure返回的Bounds在Syncfusion React PDF Viewer中高亮文本
利用Azure OCR的Bounds在Syncfusion React PDF Viewer中高亮文本
要实现这个需求,核心是将Azure OCR返回的坐标映射到Syncfusion PDF Viewer的注释系统中,通过Viewer的API添加高亮注释。以下是具体步骤和代码实现:
1. 理解坐标系统匹配
Azure Cognitive Service OCR返回的坐标以页面左上角为原点,box字段格式为x1 y1 x2 y2(左、上、右、下),与Syncfusion PDF Viewer的页面坐标体系一致。你提供的OCR数据中,页面的box值(如0 0 1224 1584)对应PDF页面的原始宽高,无需额外转换页面尺寸。
2. 解析OCR响应数据
遍历OCR返回的每个文档、页面和单词,提取每个单词的页码、坐标和文本内容:
- 注意:Azure的页码是0-based,而Syncfusion PDF Viewer的注释API使用1-based页码,需要加1转换。
3. 使用Syncfusion API添加高亮注释
通过Viewer的addAnnotation方法创建HighlightAnnotation,将每个单词的坐标映射为注释的边界。
完整代码示例
import React, { useRef, useEffect } from 'react'; import { PdfViewerComponent, Inject, HighlightAnnotation } from '@syncfusion/ej2-react-pdfviewer'; const PdfOcrHighlight = ({ ocrData, documentPath }) => { const viewerRef = useRef(null); const applyHighlights = () => { const viewerInstance = viewerRef.current?.ej2Instances; if (!viewerInstance || !ocrData) return; // 清除之前的高亮注释(可选,避免重复) viewerInstance.annotations.removeAll('Highlight'); const targetDoc = ocrData.docs[0]; targetDoc.pages.forEach(page => { const pageNumber = page.id + 1; // 转换为1-based页码 page.words.forEach(word => { // 解析单词的坐标 const [x1, y1, x2, y2] = word.box.split(' ').map(Number); // 构建高亮区域的边界 const highlightBounds = { left: x1, top: y1, width: x2 - x1, height: y2 - y1 }; // 创建高亮注释实例 const highlight = new viewerInstance.annotations.HighlightAnnotation(); highlight.pageNumber = pageNumber; highlight.bounds = highlightBounds; highlight.color = 'rgba(255, 255, 0, 0.3)'; // 黄色半透明高亮 highlight.author = 'OCR Search'; // 添加注释到Viewer viewerInstance.annotations.add(highlight); }); }); // 刷新注释渲染 viewerInstance.annotations.refresh(); }; // 当OCR数据更新时应用高亮 useEffect(() => { applyHighlights(); }, [ocrData]); return ( <div className="pdf-viewer-container"> <PdfViewerComponent ref={viewerRef} documentPath={documentPath} serviceUrl="https://ej2services.syncfusion.com/production/web-services/api/pdfviewer" style={{ height: '800px' }} > <Inject services={[HighlightAnnotation]} /> </PdfViewerComponent> </div> ); }; export default PdfOcrHighlight;
关键注意事项
- 坐标适配:如果你的PDF Viewer启用了自定义缩放,无需额外调整坐标——Syncfusion会自动基于原始页面尺寸适配注释位置。
- 短语合并:若搜索的是多词短语,可将相邻单词的坐标合并为一个大矩形,实现整段高亮,只需计算合并后的
x1(最小左坐标)、y1(最小上坐标)、x2(最大右坐标)、y2(最大下坐标)即可。 - 性能优化:当OCR返回大量单词时,可批量添加注释,减少
add方法的调用次数,提升渲染效率。
内容的提问来源于stack exchange,提问作者Sarvan
相关产品推荐
相关产品推荐

