You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在.NET(jQuery+C#)中通过点击pdf.js渲染的PDF获取字体及字号信息

点击PDF.js渲染字符获取字体与字号的实现方案

前端(JavaScript + jQuery)实现

1. 初始化PDF.js并启用文本层

PDF.js渲染时必须启用文本层(textLayer),它会将每个文本字符映射为DOM元素,同时保留原始文本的元数据(字体、字号等)。

const pdfjsLib = window['pdfjs-dist/build/pdf'];
pdfjsLib.GlobalWorkerOptions.workerSrc = 'pdf.worker.min.js'; // 替换为你的worker文件路径

async function loadPDF(pdfUrl) {
  const pdfDoc = await pdfjsLib.getDocument(pdfUrl).promise;
  const firstPage = await pdfDoc.getPage(1); // 示例加载第一页,多页场景需循环处理
  const viewport = firstPage.getViewport({ scale: 1.5 });

  const canvas = document.getElementById('pdf-canvas');
  const textLayer = document.getElementById('text-layer');

  canvas.height = viewport.height;
  canvas.width = viewport.width;

  // 渲染页面并关联文本层
  const renderTask = firstPage.render({
    canvasContext: canvas.getContext('2d'),
    viewport: viewport,
    textLayer: {
      container: textLayer,
      viewport: viewport,
      textDivs: []
    }
  });

  await renderTask.promise;
  // 绑定点击事件
  attachTextClickHandler(textLayer, firstPage);
}

2. 绑定点击事件提取字体信息

监听文本层元素的点击事件,结合PDF.js的getTextContent()方法获取页面文本元数据,匹配点击位置对应的文本项即可拿到字体和字号:

function attachTextClickHandler(textLayer, page) {
  $(textLayer).on('click', '.textLayer span', async function() {
    const clickedSpan = this;
    const textContent = await page.getTextContent();
    
    // 匹配点击元素对应的文本项(通过坐标映射)
    const spanRect = clickedSpan.getBoundingClientRect();
    const pageViewport = page.getViewport({ scale: 1 });
    const targetTextItem = textContent.items.find(item => {
      // 转换PDF坐标系到浏览器坐标系
      const itemX = item.transform[4];
      const itemY = pageViewport.height - item.transform[5];
      return itemX >= spanRect.left && itemX <= spanRect.right && itemY >= spanRect.top && itemY <= spanRect.bottom;
    });

    if (targetTextItem) {
      const fontSize = targetTextItem.height;
      const fontName = targetTextItem.fontName;
      console.log(`字体类型: ${fontName}, 字号: ${fontSize}px`);
      // 发送到后端处理
      sendToBackend({ fontName, fontSize });
    }
  });
}

// 异步发送数据到.NET后端
function sendToBackend(data) {
  $.ajax({
    url: '/api/PdfFont/SubmitFontInfo',
    method: 'POST',
    contentType: 'application/json',
    data: JSON.stringify(data),
    success: () => console.log('数据已提交'),
    error: err => console.error('提交失败:', err)
  });
}

对应的基础HTML结构:

<div style="position: relative;">
  <canvas id="pdf-canvas"></canvas>
  <div id="text-layer"></div>
</div>

文本层样式(确保可点击且不遮挡画布):

#text-layer {
  position: absolute;
  top: 0;
  left: 0;
  opacity: 0;
  pointer-events: auto;
}
#text-layer span {
  position: absolute;
  white-space: pre;
  cursor: pointer;
}

后端(C#/.NET)接口实现

创建API接口接收前端发送的字体信息,用于存储或后续业务处理:

using Microsoft.AspNetCore.Mvc;

[ApiController]
[Route("api/[controller]")]
public class PdfFontController : ControllerBase
{
    [HttpPost("SubmitFontInfo")]
    public IActionResult SubmitFontInfo([FromBody] FontDataModel model)
    {
        if (model == null)
            return BadRequest("无效请求数据");

        // 此处可添加数据库存储、日志记录等逻辑
        System.Diagnostics.Debug.WriteLine($"收到字体信息:字体={model.FontName}, 字号={model.FontSize}px");
        return Ok("数据已接收");
    }
}

// 数据传输模型
public class FontDataModel
{
    public string FontName { get; set; }
    public float FontSize { get; set; }
}

关键注意点

  • 使用PDF.js稳定版本(如v3.x),避免版本兼容问题
  • 多页PDF场景下,切换页面时需销毁旧页面的点击事件,防止内存泄漏
  • 坐标映射需注意PDF与浏览器的坐标系差异(PDF原点在左下角,浏览器在左上角)

内容的提问来源于stack exchange,提问作者g kiran

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 04:17:20