如何在.NET(jQuery+C#)中通过点击pdf.js渲染的PDF获取字体及字号信息
点击PDF.js渲染字符获取字体与字号的实现方案
前端(JavaScript + jQuery)实现
1. 初始化PDF.js并启用文本层
PDF.js渲染时必须启用文本层(textLayer),它会将每个文本字符映射为DOM元素,同时保留原始文本的元数据(字体、字号等)。
const pdfjsLib = window['pdfjs-dist/build/pdf']; pdfjsLib.GlobalWorkerOptions.workerSrc = 'pdf.worker.min.js'; // 替换为你的worker文件路径 async function loadPDF(pdfUrl) { const pdfDoc = await pdfjsLib.getDocument(pdfUrl).promise; const firstPage = await pdfDoc.getPage(1); // 示例加载第一页,多页场景需循环处理 const viewport = firstPage.getViewport({ scale: 1.5 }); const canvas = document.getElementById('pdf-canvas'); const textLayer = document.getElementById('text-layer'); canvas.height = viewport.height; canvas.width = viewport.width; // 渲染页面并关联文本层 const renderTask = firstPage.render({ canvasContext: canvas.getContext('2d'), viewport: viewport, textLayer: { container: textLayer, viewport: viewport, textDivs: [] } }); await renderTask.promise; // 绑定点击事件 attachTextClickHandler(textLayer, firstPage); }
2. 绑定点击事件提取字体信息
监听文本层元素的点击事件,结合PDF.js的getTextContent()方法获取页面文本元数据,匹配点击位置对应的文本项即可拿到字体和字号:
function attachTextClickHandler(textLayer, page) { $(textLayer).on('click', '.textLayer span', async function() { const clickedSpan = this; const textContent = await page.getTextContent(); // 匹配点击元素对应的文本项(通过坐标映射) const spanRect = clickedSpan.getBoundingClientRect(); const pageViewport = page.getViewport({ scale: 1 }); const targetTextItem = textContent.items.find(item => { // 转换PDF坐标系到浏览器坐标系 const itemX = item.transform[4]; const itemY = pageViewport.height - item.transform[5]; return itemX >= spanRect.left && itemX <= spanRect.right && itemY >= spanRect.top && itemY <= spanRect.bottom; }); if (targetTextItem) { const fontSize = targetTextItem.height; const fontName = targetTextItem.fontName; console.log(`字体类型: ${fontName}, 字号: ${fontSize}px`); // 发送到后端处理 sendToBackend({ fontName, fontSize }); } }); } // 异步发送数据到.NET后端 function sendToBackend(data) { $.ajax({ url: '/api/PdfFont/SubmitFontInfo', method: 'POST', contentType: 'application/json', data: JSON.stringify(data), success: () => console.log('数据已提交'), error: err => console.error('提交失败:', err) }); }
对应的基础HTML结构:
<div style="position: relative;"> <canvas id="pdf-canvas"></canvas> <div id="text-layer"></div> </div>
文本层样式(确保可点击且不遮挡画布):
#text-layer { position: absolute; top: 0; left: 0; opacity: 0; pointer-events: auto; } #text-layer span { position: absolute; white-space: pre; cursor: pointer; }
后端(C#/.NET)接口实现
创建API接口接收前端发送的字体信息,用于存储或后续业务处理:
using Microsoft.AspNetCore.Mvc; [ApiController] [Route("api/[controller]")] public class PdfFontController : ControllerBase { [HttpPost("SubmitFontInfo")] public IActionResult SubmitFontInfo([FromBody] FontDataModel model) { if (model == null) return BadRequest("无效请求数据"); // 此处可添加数据库存储、日志记录等逻辑 System.Diagnostics.Debug.WriteLine($"收到字体信息:字体={model.FontName}, 字号={model.FontSize}px"); return Ok("数据已接收"); } } // 数据传输模型 public class FontDataModel { public string FontName { get; set; } public float FontSize { get; set; } }
关键注意点
- 使用PDF.js稳定版本(如v3.x),避免版本兼容问题
- 多页PDF场景下,切换页面时需销毁旧页面的点击事件,防止内存泄漏
- 坐标映射需注意PDF与浏览器的坐标系差异(PDF原点在左下角,浏览器在左上角)
内容的提问来源于stack exchange,提问作者g kiran
相关产品推荐
相关产品推荐

