Java正则表达式提取肯尼亚证件关键数据的正确实现咨询
Fixing Kenyan ID Card OCR Data Extraction with Java Regex
Hey there! Let's get your regex sorted to pull out those key fields from the Kenyan ID OCR text properly. The main issue with your current code is that it's matching numbers too broadly instead of targeting each field by its specific label in the text. Here's a tailored solution that works line-by-line to capture exactly what you need:
Modified Code with Targeted Regex
static HashMap<String, String> interpretText(String ocrText) { HashMap<String, String> result = new HashMap<String, String>(); result.put("Text", ocrText); // Pre-compile patterns once outside the loop for better performance Pattern idSerialPattern = Pattern.compile("senne wnecs:\\s*(\\d+)\\s*\\w+\\s*(\\d+)"); Pattern fullNamePattern = Pattern.compile("FULL NAMES\\s+(.*)"); Pattern dobPattern = Pattern.compile("DATE OF BIRTH\\s+(\\d{2}\\.\\s*\\d{2}\\.\\s*\\d{4})"); for (String line : ocrText.split("\\r?\\n")) { String cleanedLine = line.trim(); if (cleanedLine.isEmpty()) continue; // Capture ID Number and Serial Number from the relevant line Matcher idSerialMatcher = idSerialPattern.matcher(cleanedLine); if (idSerialMatcher.find()) { String idNumber = idSerialMatcher.group(1).trim(); String serialNumber = idSerialMatcher.group(2).trim(); result.put("IDNumber", idNumber); result.put("SerialNumber", serialNumber); System.out.printf("IDNumber: %s%nSerialNumber: %s%n", idNumber, serialNumber); continue; // Move to next line since we found what we need here } // Capture Full Name Matcher nameMatcher = fullNamePattern.matcher(cleanedLine); if (nameMatcher.find()) { String fullName = nameMatcher.group(1).trim(); result.put("Name", fullName); System.out.println("Name: " + fullName); continue; } // Capture Date of Birth Matcher dobMatcher = dobPattern.matcher(cleanedLine); if (dobMatcher.find()) { String dateOfBirth = dobMatcher.group(1).trim(); result.put("DateOfBirth", dateOfBirth); System.out.println("DateOfBirth: " + dateOfBirth); continue; } } return result; }
How This Works
Let's break down each regex pattern to explain what it does:
ID & Serial Number Pattern:
senne wnecs:\\s*(\\d+)\\s*\\w+\\s*(\\d+)- Matches the line containing your ID and serial numbers (
sennc wnecs: 23085129 aShl e 31662252) \\s*handles any extra spaces between elements, while(\\d+)captures the numeric values for ID and Serial separately.
- Matches the line containing your ID and serial numbers (
Full Name Pattern:
FULL NAMES\\s+(.*)- Targets the line starting with
FULL NAMESand captures everything after the label (your full name).
- Targets the line starting with
Date of Birth Pattern:
DATE OF BIRTH\\s+(\\d{2}\\.\\s*\\d{2}\\.\\s*\\d{4})- Matches the date format in your OCR text (e.g.,
25. 10. 1992), accounting for optional spaces after the dots with\\s*.
- Matches the date format in your OCR text (e.g.,
Extra Tips
- If the OCR has slight spelling variations for labels (like
senne wnecsbeing a misread ofSerial No), you can make the regex more flexible. For example, use(?i)serial\\s*no\\s*:to match case-insensitively and handle spacing. - Pre-compiling patterns outside the loop improves efficiency, especially if you're processing multiple OCR texts.
内容的提问来源于stack exchange,提问作者Cheruiyot Felix
相关产品推荐
相关产品推荐

