You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

HashMap打印出现异常数值原因及词频统计程序开发疑问

Q1: Why Am I Seeing a Strange Number When Printing My HashMap Values?

Ah, that's a common gotcha! Let's walk through the most likely culprits:

  • You're printing the object itself, not the actual value: If you're doing something like System.out.println(myHashMap) or System.out.println(entry) instead of entry.getValue(), Java's default toString() for HashMap entries spits out something like key@123abc—that hex number is the object's hash code. Always explicitly grab the value when printing.
  • You accidentally stored the hash code instead of the count: Double-check your count logic. If you wrote wordCount.put(word, word.hashCode()) instead of incrementing the count (like wordCount.put(word, wordCount.getOrDefault(word, 0) + 1)), you're storing a random-ish number instead of how many times the word appears.
  • Integer overflow (rare but possible): If you're processing a massive text file, the count could exceed the max value of Integer (2,147,483,647). When that happens, the number wraps around to negative or weird values—switch to Long for your counts to fix this.
  • Bad parsing from the file: If your common words file has numbers mixed in, and you're accidentally parsing those into your count map, that could cause strange numbers. But since this is a word count program, this is less likely unless your delimiter is letting numbers slip through.
Q2: Completing the Common Words HashMap Setup

Looking at your code snippet, here's how to finish populating that commonHashMap properly, plus some tips to avoid headaches later:

First, let's fix the loop to read each word correctly and mark it as common (we just need a way to check if a word is in the list, so storing true as the value works perfectly):

try { 
    // Populate commonHashMap
    Scanner sc = new Scanner(new File("commonwords.txt")); 
    sc.useDelimiter("[^a-zA-Z']"); 
    String str; 
    while (sc.hasNext()) { 
        str = sc.next().toLowerCase(); // Convert to lowercase to avoid case issues
        // Skip empty strings that might come from punctuation/delimiter quirks
        if (!str.isEmpty()) {
            commonHashMap.put(str, true); // Mark this word as common
        }
    }
    sc.close(); // Don't forget to clean up the scanner!
} catch (FileNotFoundException e) {
    // Better to handle this gracefully than just print the stack trace
    System.err.println("Couldn't find commonwords.txt: " + e.getMessage());
}

A few key tips for your word count program:

  • Case insensitivity is non-negotiable: Convert every word to lowercase (or uppercase) when adding to both your common words map and your count map. Otherwise, "Hello" and "hello" will be counted separately, which is almost never what you want.
  • Skip empty strings: Your delimiter pattern might generate empty strings (even with apostrophes allowed), so adding that check saves you from adding garbage to your map.
  • Use try-with-resources for safety: Instead of manually closing the scanner, use a try-with-resources block to auto-close it. This prevents resource leaks if an error occurs mid-read:
    try (Scanner sc = new Scanner(new File("commonwords.txt"))) {
        // Your loop here
    } catch (FileNotFoundException e) {
        // Handle the error
    }
    
  • Filtering common words: When you're done counting words, iterate through your count map and only keep entries where the word is not in commonHashMap and the count is above the user's threshold. That way, you exclude those overused words from your results.

内容的提问来源于stack exchange,提问作者Hüseyin Eren Güler

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 03:33:59