You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Java跨文件字符对比求助:字符差异统计功能异常排查

Fixing Your Character Error Rate (CER) Calculation in Java

Let's walk through fixing your CER code—right now it has a few critical bugs that are keeping it from correctly identifying and counting differing characters, just like in your example where you expect 3 differences between "my name" and "mi n@mes".

First, let's break down the issues in your original code:

  • The end variable is initialized to 0, so the while loop never runs at all—meaning it never checks individual characters.
  • You accidentally compare originalteks.charAt(start) to itself instead of extractteks.charAt(start)—so it's never actually checking the extracted text against the original.
  • The logic only accounts for length differences but doesn't handle character-by-character mismatches when lengths match (or even when they don't, beyond the initial length delta).

Here's the corrected code that works as expected:

public float getCER(String originalTeks, String extractTeks) {
    int differentChar = 0;
    int originalLength = originalTeks.length();
    int extractLength = extractTeks.length();

    // Account for extra/missing characters from length differences
    differentChar += Math.abs(originalLength - extractLength);

    // Compare characters up to the length of the shorter string
    int minLength = Math.min(originalLength, extractLength);
    for (int i = 0; i < minLength; i++) {
        if (originalTeks.charAt(i) != extractTeks.charAt(i)) {
            differentChar++;
        }
    }

    // Calculate final CER: total differences divided by original text length
    return (float) differentChar / originalLength;
}

Let's test this with your example:

  • Original text: "my name" (length 7)
  • Extracted text: "mi n@mes" (length 8)
  1. Length difference is 8-7=1, so differentChar starts at 1.
  2. Comparing each character up to 7 positions:
    • Position 1: y vs i → mismatch, count becomes 2
    • Position 4: a vs @ → mismatch, count becomes 3
  3. Total differences: 3, which exactly matches your expected result.
  4. The final CER would be 3/7 ≈ 0.4286.

Key improvements made:

  • Renamed variables to follow Java camelCase conventions for readability
  • Fixed the character comparison logic to actually check original vs extracted text
  • Combined length differences with mismatched character counts for accurate total error count
  • Swapped the broken while loop for a straightforward for loop to iterate through characters clearly

内容的提问来源于stack exchange,提问作者michael

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 03:09:49