本地化数字解析时非法输入未触发ParseException的问题与处理
Great question! This is a common gotcha with Java's DecimalFormat—let's break down why this behavior exists and how to enforce strict parsing for your use case.
Why Doesn't DecimalFormat.parse() Fail Fast?
The partial parsing behavior you're seeing is intentional design from the JDK team. Here's the rationale:
- It's built to handle real-world, non-strict input scenarios. For example, if a user enters "123.45$" or "500 元", many applications want to extract the valid number part instead of throwing an immediate error.
- The
parse()method is documented to read from the start of the string until it hits the first unrecognizable character, then return the successfully parsed portion. It usesParsePosition(which you can explicitly pass) to indicate where parsing stopped—this gives developers control over whether to accept partial results or enforce strictness.
In your example, "123456Test99" gets parsed up to "123456" because 'T' is not a valid character for the Ukrainian locale's number format, so parsing stops there instead of throwing an exception.
How to Enforce Strict Parsing
To ensure the entire input string is a valid number (no extra characters), you have two reliable approaches:
1. Use ParsePosition to Verify Full Parsing (Recommended)
This is the official, robust way to check if the entire string was parsed. It handles all edge cases supported by DecimalFormat (like grouping separators, scientific notation, and locale-specific symbols) without manual character checks.
import java.math.BigDecimal; import java.text.DecimalFormat; import java.text.ParseException; import java.text.ParsePosition; import java.util.Locale; public class StrictNumberParser { public static BigDecimal parseStrictly(DecimalFormat format, String input) throws ParseException { if (input == null || input.trim().isEmpty()) { return null; } input = input.trim(); ParsePosition pos = new ParsePosition(0); Object parsedValue = format.parse(input, pos); // Check if parsing consumed the entire string if (pos.getIndex() != input.length()) { throw new ParseException( String.format("Invalid character at position %d", pos.getIndex()), pos.getIndex() ); } // Since we set setParseBigDecimal(true), cast safely return (BigDecimal) parsedValue; } public static void main(String[] args) throws ParseException { Locale ukLocale = Locale.forLanguageTag("uk-UK"); DecimalFormat decimalFormat = (DecimalFormat) NumberFormat.getInstance(ukLocale); decimalFormat.setParseBigDecimal(true); // Valid case - works as expected System.out.println(parseStrictly(decimalFormat, "123456,99")); // Output: 123456.99 // Invalid case - throws ParseException try { parseStrictly(decimalFormat, "123456Test99"); } catch (ParseException e) { System.out.println("Expected error: " + e.getMessage()); // Output: Invalid character at position 6 } } }
2. Enhanced Character Validation
If you prefer to pre-validate characters (as you tried), you need to cover all valid symbols supported by DecimalFormat (your original code missed plus signs, scientific notation exponents, etc.). Here's an improved version:
import java.math.BigDecimal; import java.text.DecimalFormat; import java.text.DecimalFormatSymbols; import java.text.ParseException; import java.util.Locale; public class StrictCharValidator { public static BigDecimal processNumber(DecimalFormat format, String input) throws ParseException { if (input == null || input.trim().isEmpty()) { return null; } input = input.trim(); DecimalFormatSymbols symbols = format.getDecimalFormatSymbols(); final char groupingSep = symbols.getGroupingSeparator(); final char decimalSep = symbols.getDecimalSeparator(); final char minusSign = symbols.getMinusSign(); final char plusSign = symbols.getPlusSign(); final char expUpper = 'E'; final char expLower = 'e'; for (int i = 0; i < input.length(); i++) { char ch = input.charAt(i); boolean isAllowed = Character.isDigit(ch) || ch == decimalSep || ch == groupingSep || ch == minusSign || ch == plusSign || ch == expUpper || ch == expLower; if (!isAllowed) { throw new ParseException( String.format("Invalid character '%c' at position %d", ch, i), i ); } // Ensure signs only appear at start or after exponent if ((ch == minusSign || ch == plusSign) && i != 0) { char prevChar = input.charAt(i - 1); if (prevChar != expUpper && prevChar != expLower) { throw new ParseException( String.format("Sign '%c' can only be at start or after exponent (position %d)", ch, i), i ); } } } // Final parse check to catch invalid formats (e.g., multiple decimals) ParsePosition pos = new ParsePosition(0); BigDecimal parsed = (BigDecimal) format.parse(input, pos); if (pos.getIndex() != input.length()) { throw new ParseException("Failed to parse entire string", pos.getIndex()); } return parsed; } public static void main(String[] args) throws ParseException { Locale ukLocale = Locale.forLanguageTag("uk-UK"); DecimalFormat decimalFormat = (DecimalFormat) NumberFormat.getInstance(ukLocale); decimalFormat.setParseBigDecimal(true); System.out.println(processNumber(decimalFormat, "123456,99")); // Output: 123456.99 try { processNumber(decimalFormat, "123456Test99"); } catch (ParseException e) { System.out.println("Expected error: " + e.getMessage()); // Output: Invalid character 'T' at position 6 } } }
Key Takeaway
The first approach (using ParsePosition) is preferred because it leverages DecimalFormat's built-in logic and avoids missing edge cases (like scientific notation or locale-specific symbols you might forget to validate manually).
内容的提问来源于stack exchange,提问作者Vlad

