You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

咨询:如何用Java读取扫描PDF表单填写框坐标及数据适配JS表单

解决步骤

一、用Java识别扫描PDF中填写框的精准坐标

扫描PDF本质是内嵌图像的PDF,需先提取图像再做图像分析:

  1. 提取PDF中的扫描图像
    用Apache PDFBox库加载PDF,提取页面内的图像:

    PDDocument document = PDDocument.load(new File("scanned-form.pdf"));
    PDPage page = document.getPage(0); // 默认处理第一页,可按需遍历多页
    PDResources resources = page.getResources();
    for (COSName name : resources.getXObjectNames()) {
        PDXObject xobject = resources.getXObject(name);
        if (xobject instanceof PDImageXObject) {
            PDImageXObject image = (PDImageXObject) xobject;
            BufferedImage bufferedImage = image.getImage();
            // 后续基于该BufferedImage处理
        }
    }
    document.close();
    
  2. 图像预处理与轮廓检测
    通过OpenCV对图像做灰度化、二值化,突出填写框的矩形轮廓:

    // 将BufferedImage转换为OpenCV的Mat对象(需自行实现转换工具方法)
    Mat mat = BufferedImage2Mat(bufferedImage);
    // 灰度化
    Imgproc.cvtColor(mat, mat, Imgproc.COLOR_BGR2GRAY);
    // 二值化(阈值可根据表单实际情况调整)
    Imgproc.threshold(mat, mat, 127, 255, Imgproc.THRESH_BINARY_INV);
    // 提取轮廓
    List<MatOfPoint> contours = new ArrayList<>();
    Mat hierarchy = new Mat();
    Imgproc.findContours(mat, contours, hierarchy, Imgproc.RETR_EXTERNAL, Imgproc.CHAIN_APPROX_SIMPLE);
    
  3. 筛选填写框并转换坐标
    遍历轮廓筛选符合填写框尺寸的矩形,同时将图像坐标转换为PDF标准坐标(PDF原点在左下角,图像原点在左上角):

    // 定义存储填写框信息的实体类
    class FormField {
        private float x;
        private float y;
        private float width;
        private float height;
        private int pageNum;
        private String text; // 后续存储识别的手写内容
        // 构造器、getter/setter方法
    }
    
    List<FormField> fields = new ArrayList<>();
    PDPageMediaBox mediaBox = page.getMediaBox();
    float pdfPageHeight = mediaBox.getHeight();
    float scaleX = mediaBox.getWidth() / bufferedImage.getWidth();
    float scaleY = pdfPageHeight / bufferedImage.getHeight();
    
    for (MatOfPoint contour : contours) {
        Rect rect = Imgproc.boundingRect(contour);
        // 根据表单实际尺寸过滤无效轮廓
        if (rect.width > 40 && rect.height > 15 && rect.width < 600 && rect.height < 120) {
            float pdfX = rect.x * scaleX;
            // 转换y坐标:PDF页面高度 - 图像y坐标 - 矩形高度
            float pdfY = pdfPageHeight - (rect.y + rect.height) * scaleY;
            fields.add(new FormField(pdfX, pdfY, rect.width * scaleX, rect.height * scaleY, 0, ""));
        }
    }
    

二、提取填写框内的手写内容

用Tess4J(Tesseract的Java封装)对每个填写框区域做OCR识别:

  1. 裁剪填写框区域
    根据识别出的图像坐标,从原图像中裁剪出目标区域:

    BufferedImage fieldImage = bufferedImage.getSubimage(rect.x, rect.y, rect.width, rect.height);
    
  2. OCR识别手写文本
    初始化Tesseract引擎,识别裁剪后的图像内容:

    ITesseract tesseract = new Tesseract();
    tesseract.setDatapath("tessdata"); // 指向Tesseract语言包存放目录
    tesseract.setLanguage("eng"); // 中文手写可切换为"chi_sim"
    String text = tesseract.doOCR(fieldImage);
    // 清理识别结果,去除多余空白字符
    text = text.trim().replaceAll("\\s+", " ");
    // 将文本绑定到对应的FormField对象
    field.setText(text);
    
  3. 生成结构化数据
    将包含坐标、文本的FormField列表序列化为JSON,供前端调用。

三、同步到前端JS表单

  1. 后端接口传递数据
    编写REST接口,将结构化的表单识别数据返回给前端。

  2. 表单填充逻辑
    前端可通过两种方式匹配填充:

    • 字段名匹配(推荐):提前为每个PDF填写框和前端表单输入框设置统一的字段名,直接根据字段名映射填充:
      const formData = await fetch('/api/form-recognize').then(res => res.json());
      formData.forEach(item => {
          const input = document.querySelector(`input[name="${item.fieldName}"]`);
          if (input) input.value = item.text;
      });
      
    • 坐标映射:计算PDF页面与前端表单的缩放比例,将PDF坐标转换为页面坐标,定位到对应输入框填充(适合无明确字段名的场景)。
  3. Chrome前端原生处理(可选)
    若无需后端,可直接用Chrome支持的pdf.js加载PDF,结合opencv.js做轮廓检测、Tesseract.js做OCR,全程在前端完成识别与填充,流程与后端Java逻辑一致。

内容的提问来源于stack exchange,提问作者Roberto33911

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.05 21:37:05