You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何借助Azure返回的Bounds在Syncfusion React PDF Viewer中高亮文本

利用Azure OCR的Bounds在Syncfusion React PDF Viewer中高亮文本

要实现这个需求,核心是将Azure OCR返回的坐标映射到Syncfusion PDF Viewer的注释系统中,通过Viewer的API添加高亮注释。以下是具体步骤和代码实现:

1. 理解坐标系统匹配

Azure Cognitive Service OCR返回的坐标以页面左上角为原点,box字段格式为x1 y1 x2 y2(左、上、右、下),与Syncfusion PDF Viewer的页面坐标体系一致。你提供的OCR数据中,页面的box值(如0 0 1224 1584)对应PDF页面的原始宽高,无需额外转换页面尺寸。

2. 解析OCR响应数据

遍历OCR返回的每个文档、页面和单词,提取每个单词的页码、坐标和文本内容:

  • 注意:Azure的页码是0-based,而Syncfusion PDF Viewer的注释API使用1-based页码,需要加1转换。

3. 使用Syncfusion API添加高亮注释

通过Viewer的addAnnotation方法创建HighlightAnnotation,将每个单词的坐标映射为注释的边界。

完整代码示例

import React, { useRef, useEffect } from 'react';
import { PdfViewerComponent, Inject, HighlightAnnotation } from '@syncfusion/ej2-react-pdfviewer';

const PdfOcrHighlight = ({ ocrData, documentPath }) => {
  const viewerRef = useRef(null);

  const applyHighlights = () => {
    const viewerInstance = viewerRef.current?.ej2Instances;
    if (!viewerInstance || !ocrData) return;

    // 清除之前的高亮注释(可选,避免重复)
    viewerInstance.annotations.removeAll('Highlight');

    const targetDoc = ocrData.docs[0];
    targetDoc.pages.forEach(page => {
      const pageNumber = page.id + 1; // 转换为1-based页码
      page.words.forEach(word => {
        // 解析单词的坐标
        const [x1, y1, x2, y2] = word.box.split(' ').map(Number);
        // 构建高亮区域的边界
        const highlightBounds = {
          left: x1,
          top: y1,
          width: x2 - x1,
          height: y2 - y1
        };

        // 创建高亮注释实例
        const highlight = new viewerInstance.annotations.HighlightAnnotation();
        highlight.pageNumber = pageNumber;
        highlight.bounds = highlightBounds;
        highlight.color = 'rgba(255, 255, 0, 0.3)'; // 黄色半透明高亮
        highlight.author = 'OCR Search';

        // 添加注释到Viewer
        viewerInstance.annotations.add(highlight);
      });
    });

    // 刷新注释渲染
    viewerInstance.annotations.refresh();
  };

  // 当OCR数据更新时应用高亮
  useEffect(() => {
    applyHighlights();
  }, [ocrData]);

  return (
    <div className="pdf-viewer-container">
      <PdfViewerComponent
        ref={viewerRef}
        documentPath={documentPath}
        serviceUrl="https://ej2services.syncfusion.com/production/web-services/api/pdfviewer"
        style={{ height: '800px' }}
      >
        <Inject services={[HighlightAnnotation]} />
      </PdfViewerComponent>
    </div>
  );
};

export default PdfOcrHighlight;

关键注意事项

  • 坐标适配:如果你的PDF Viewer启用了自定义缩放,无需额外调整坐标——Syncfusion会自动基于原始页面尺寸适配注释位置。
  • 短语合并:若搜索的是多词短语,可将相邻单词的坐标合并为一个大矩形,实现整段高亮,只需计算合并后的x1(最小左坐标)、y1(最小上坐标)、x2(最大右坐标)、y2(最大下坐标)即可。
  • 性能优化:当OCR返回大量单词时,可批量添加注释,减少add方法的调用次数,提升渲染效率。

内容的提问来源于stack exchange,提问作者Sarvan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.29 13:05:24