You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Zxing与Apache PDFBox提取PDF条形码时坐标超出范围的问题及解决方案咨询

Zxing与Apache PDFBox提取PDF条形码时坐标超出范围的问题及解决方案咨询

大家好,我现在遇到一个棘手的问题,想请教下各位大佬。需求是提取PDF文件里的条形码列表,包括每个条码的文本、格式,以及它们在PDF页面上的准确位置。目前我用的是开源的Apache PDFBox和Zxing Java库,项目要求不能用付费工具,所以只能在这两个库的基础上调整。

现在的情况是,条形码的文本和格式都能正常提取,但提取到的坐标值完全超出了PDF的正常范围,比如像Point: (4791.0, 589.5)和Point: (4791.0, 1001.5)这样的数值,根本没法对应到PDF页面的实际位置上。

先给大家看下当前代码输出的结果:

当前代码输出

Barcode text: https://qr.aa/1133
Barcode format: QR_CODE
Point: (4689.0, 3745.0)
Point: (4689.0, 3569.0)
Point: (4865.0, 3569.0)
Point: (4841.0, 3721.0)

Barcode text: 28852
Barcode format: CODE_128
Point: (4791.0, 589.5)
Point: (4791.0, 1001.5)

Barcode text: 08179018
Barcode format: UPC_E
Point: (1500.5, 438.0)
Point: (906.5, 438.0)

Barcode text: 147089001
Barcode format: CODE_128
Point: (4867.0, 1522.5)
Point: (4867.0, 2072.5)

Barcode text: OPS
Barcode format: CODE_128
Point: (4867.0, 2603.5)
Point: (4867.0, 2947.5)

然后是我现在用的代码(刚才检查的时候发现里面有个嵌套循环的笔误,重复遍历了ResultPoint,已经修正后贴出来了):

public static void getBarcodePosition(String filename) throws IOException, NotFoundException {
    PDDocument document = Loader.loadPDF(new File(filename));

    PDFRenderer pdfRenderer = new PDFRenderer(document);

    for (int page = 0; page < document.getNumberOfPages(); ++page) {
        BufferedImage image = pdfRenderer.renderImageWithDPI(page, 600, ImageType.RGB);
        LuminanceSource source = new BufferedImageLuminanceSource(image);
        BinaryBitmap bitmap = new BinaryBitmap(new HybridBinarizer(source));

        Hashtable<DecodeHintType, Object> hints = new Hashtable<>();
        hints.put(DecodeHintType.TRY_HARDER, Boolean.TRUE);
        GenericMultipleBarcodeReader reader = new GenericMultipleBarcodeReader(new MultiFormatReader());
        Result[] results = reader.decodeMultiple(bitmap, hints);

        for (Result result : results) {
            System.out.println("Barcode text: " + result.getText());
            System.out.println("Barcode format: " + result.getBarcodeFormat());
            // 修正了原代码中重复嵌套的ResultPoint遍历循环
            for (ResultPoint point : result.getResultPoints()) {
                System.out.println("Point: (" + point.getX() + ", " + point.getY() + ")");
            }
            System.out.println("                                                        ");
        }
    }
    document.close();
}

想请教下大家,有没有办法修改这段代码来获取正确的PDF坐标?或者有没有其他靠谱的替代实现可以参考?非常感谢!

备注:内容来源于stack exchange,提问作者Chiranjib

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.13 19:35:26