使用Zxing与Apache PDFBox提取PDF条形码时坐标超出范围的问题及解决方案咨询
大家好,我现在遇到一个棘手的问题,想请教下各位大佬。需求是提取PDF文件里的条形码列表,包括每个条码的文本、格式,以及它们在PDF页面上的准确位置。目前我用的是开源的Apache PDFBox和Zxing Java库,项目要求不能用付费工具,所以只能在这两个库的基础上调整。
现在的情况是,条形码的文本和格式都能正常提取,但提取到的坐标值完全超出了PDF的正常范围,比如像Point: (4791.0, 589.5)和Point: (4791.0, 1001.5)这样的数值,根本没法对应到PDF页面的实际位置上。
先给大家看下当前代码输出的结果:
当前代码输出
Barcode text: https://qr.aa/1133
Barcode format: QR_CODE
Point: (4689.0, 3745.0)
Point: (4689.0, 3569.0)
Point: (4865.0, 3569.0)
Point: (4841.0, 3721.0)
Barcode text: 28852
Barcode format: CODE_128
Point: (4791.0, 589.5)
Point: (4791.0, 1001.5)
Barcode text: 08179018
Barcode format: UPC_E
Point: (1500.5, 438.0)
Point: (906.5, 438.0)
Barcode text: 147089001
Barcode format: CODE_128
Point: (4867.0, 1522.5)
Point: (4867.0, 2072.5)
Barcode text: OPS
Barcode format: CODE_128
Point: (4867.0, 2603.5)
Point: (4867.0, 2947.5)
然后是我现在用的代码(刚才检查的时候发现里面有个嵌套循环的笔误,重复遍历了ResultPoint,已经修正后贴出来了):
public static void getBarcodePosition(String filename) throws IOException, NotFoundException { PDDocument document = Loader.loadPDF(new File(filename)); PDFRenderer pdfRenderer = new PDFRenderer(document); for (int page = 0; page < document.getNumberOfPages(); ++page) { BufferedImage image = pdfRenderer.renderImageWithDPI(page, 600, ImageType.RGB); LuminanceSource source = new BufferedImageLuminanceSource(image); BinaryBitmap bitmap = new BinaryBitmap(new HybridBinarizer(source)); Hashtable<DecodeHintType, Object> hints = new Hashtable<>(); hints.put(DecodeHintType.TRY_HARDER, Boolean.TRUE); GenericMultipleBarcodeReader reader = new GenericMultipleBarcodeReader(new MultiFormatReader()); Result[] results = reader.decodeMultiple(bitmap, hints); for (Result result : results) { System.out.println("Barcode text: " + result.getText()); System.out.println("Barcode format: " + result.getBarcodeFormat()); // 修正了原代码中重复嵌套的ResultPoint遍历循环 for (ResultPoint point : result.getResultPoints()) { System.out.println("Point: (" + point.getX() + ", " + point.getY() + ")"); } System.out.println(" "); } } document.close(); }
想请教下大家,有没有办法修改这段代码来获取正确的PDF坐标?或者有没有其他靠谱的替代实现可以参考?非常感谢!
备注:内容来源于stack exchange,提问作者Chiranjib

