Java跨文件字符对比求助:字符差异统计功能异常排查
Fixing Your Character Error Rate (CER) Calculation in Java
Let's walk through fixing your CER code—right now it has a few critical bugs that are keeping it from correctly identifying and counting differing characters, just like in your example where you expect 3 differences between "my name" and "mi n@mes".
First, let's break down the issues in your original code:
- The
endvariable is initialized to 0, so thewhileloop never runs at all—meaning it never checks individual characters. - You accidentally compare
originalteks.charAt(start)to itself instead ofextractteks.charAt(start)—so it's never actually checking the extracted text against the original. - The logic only accounts for length differences but doesn't handle character-by-character mismatches when lengths match (or even when they don't, beyond the initial length delta).
Here's the corrected code that works as expected:
public float getCER(String originalTeks, String extractTeks) { int differentChar = 0; int originalLength = originalTeks.length(); int extractLength = extractTeks.length(); // Account for extra/missing characters from length differences differentChar += Math.abs(originalLength - extractLength); // Compare characters up to the length of the shorter string int minLength = Math.min(originalLength, extractLength); for (int i = 0; i < minLength; i++) { if (originalTeks.charAt(i) != extractTeks.charAt(i)) { differentChar++; } } // Calculate final CER: total differences divided by original text length return (float) differentChar / originalLength; }
Let's test this with your example:
- Original text:
"my name"(length 7) - Extracted text:
"mi n@mes"(length 8)
- Length difference is
8-7=1, sodifferentCharstarts at 1. - Comparing each character up to 7 positions:
- Position 1:
yvsi→ mismatch, count becomes 2 - Position 4:
avs@→ mismatch, count becomes 3
- Position 1:
- Total differences: 3, which exactly matches your expected result.
- The final CER would be
3/7 ≈ 0.4286.
Key improvements made:
- Renamed variables to follow Java camelCase conventions for readability
- Fixed the character comparison logic to actually check original vs extracted text
- Combined length differences with mismatched character counts for accurate total error count
- Swapped the broken
whileloop for a straightforwardforloop to iterate through characters clearly
内容的提问来源于stack exchange,提问作者michael
相关产品推荐
相关产品推荐

