You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于给定词库的缩写展开程序提交失败排查求助

Troubleshooting Your Blog Abbreviation Expander

Let's break down the potential issues in your code that might be causing failures on the test server. I've identified several edge-case problems that aren't covered in your example input:

1. Incomplete Punctuation Handling

Your current code only handles a single trailing punctuation mark (like ? or !) from the abbreviation, but fails when multiple punctuation marks are present (e.g., bst?!). Instead of just taking the last character, you need to extract all trailing punctuation from the abbreviation, match the core abbreviation body to the lexicon, then reattach all punctuation to the matched word.

Fix:

Add a helper method to separate the abbreviation body from trailing punctuation, then use only the body for matching:

private static boolean isPunctuation(char c) {
    return c == '!' || c == '?' || c == ',' || c == ';';
}

// Inside inputExpander, for each abbrToken:
StringBuilder punctuation = new StringBuilder();
int lastNonPunctIdx = abbrToken.length() - 1;
while (lastNonPunctIdx >= 0 && isPunctuation(abbrToken.charAt(lastNonPunctIdx))) {
    punctuation.insert(0, abbrToken.charAt(lastNonPunctIdx));
    lastNonPunctIdx--;
}
String abbrBody = lastNonPunctIdx >= 0 ? abbrToken.substring(0, lastNonPunctIdx + 1) : "";
char[] abbrBodyChars = abbrBody.toCharArray();

Then use abbrBodyChars instead of char_abbr_token for the ifcontains check. After finding the best match, append the punctuation string to it.

2. Incorrect Matching of Abbreviations with Embedded Punctuation

Your ifcontains method treats any punctuation in the abbreviation as a match, even if it's not at the end. For example, an abbreviation like l?ke would incorrectly match like because the method flags the ? as a valid match instead of ignoring it. By separating punctuation first (as above), you avoid this issue entirely—you only match the core abbreviation body against lexicon words.

Fix:

Update the ifcontains method to only process the abbreviation body (no punctuation), and remove the logic that handles punctuation inside the abbreviation:

public static boolean ifcontains(char[] char_text_token, char[] char_abbr_body) {
    int j = 0;
    for (int i = 0; i < char_abbr_body.length; ++i) {
        boolean found = false;
        for (; j < char_text_token.length; ++j) {
            if (char_abbr_body[i] == char_text_token[j]) {
                found = true;
                j++;
                break;
            }
        }
        if (!found) {
            return false;
        }
    }
    return true;
}

3. Length Comparison Logic is Fragile

Your current code adjusts the stored word's length by subtracting 1 if it has punctuation, but this breaks if there are multiple punctuation marks. By separating the punctuation from the matched word, you can compare the raw lengths of the lexicon words directly (since punctuation doesn't count towards the "shortest word" rule).

Fix:

When comparing candidate matches, use the original lexicon word's length (without added punctuation). After selecting the best match, append the saved punctuation string to get the final output token.

4. Empty Abbreviation Body Handling

If an abbreviation is made up entirely of punctuation (e.g., !!!), your code should return the original abbreviation instead of trying to match it to the lexicon. The fix from point 1 already handles this by checking if abbrBody is empty.

Revised InputExpander Method Snippet

Here's how the updated inputExpander would look with these fixes:

public static String[] inputExpander(String[] text, String[] abbr) {
    String[] output = new String[abbr.length];
    
    for (int i = 0; i < abbr.length; ++i) {
        String abbrToken = abbr[i];
        // Separate abbreviation body and trailing punctuation
        StringBuilder punctuation = new StringBuilder();
        int lastNonPunctIdx = abbrToken.length() - 1;
        while (lastNonPunctIdx >= 0 && isPunctuation(abbrToken.charAt(lastNonPunctIdx))) {
            punctuation.insert(0, abbrToken.charAt(lastNonPunctIdx));
            lastNonPunctIdx--;
        }
        String abbrBody = lastNonPunctIdx >= 0 ? abbrToken.substring(0, lastNonPunctIdx + 1) : "";
        char[] abbrBodyChars = abbrBody.toCharArray();
        
        String bestMatch = null;
        int minLength = Integer.MAX_VALUE;
        boolean hasMultipleMatches = false;
        
        for (String textToken : text) {
            char[] textChars = textToken.toCharArray();
            if (ifcontains(textChars, abbrBodyChars)) {
                int tokenLength = textToken.length();
                if (bestMatch == null) {
                    bestMatch = textToken;
                    minLength = tokenLength;
                } else if (tokenLength < minLength) {
                    bestMatch = textToken;
                    minLength = tokenLength;
                    hasMultipleMatches = false;
                } else if (tokenLength == minLength) {
                    hasMultipleMatches = true;
                }
            }
        }
        
        // Determine final output token
        if (hasMultipleMatches || bestMatch == null) {
            output[i] = abbrToken;
        } else {
            output[i] = bestMatch + punctuation.toString();
        }
    }
    return output;
}

These changes should handle edge cases like multiple punctuation marks, empty abbreviation bodies, and ensure correct length comparison for matches.

内容的提问来源于stack exchange,提问作者user9777638

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 10:08:02