You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

NLP项目中使用Python Counter()统计文档多次出现单词的技术问询

筛选Counter中出现次数不止一次的单词

Hey there! Let's tackle this for your NLP project. You're already using collections.Counter() to get word frequencies, and filtering out words that appear more than once is totally straightforward with a bit of Python code.

First, a quick heads-up: I noticed in your shared Counter output, the entry 'evident': , is missing a count value. That's probably a typo when you shared it, just make sure your actual Counter object has valid key-value pairs (each word mapped to an integer count) before running the code below.

Solution Code

Here's how you can filter your results exactly as you need:

from collections import Counter

# Replace this with your actual Counter object (fixed the 'evident' count as an example)
word_counts = Counter({
    'due': 23, 'support': 20, 'ATM': 16, 'come': 12, 'case': 11,
    'Sallu': 10, 'tough,': 9, 'team': 8, 'evident': 7, 'likely': 6,
    'rupee': 4, 'depreciated': 2, 'senior': 1, 'neutral': 1, 'told': 1,
    'tour Russia’s': 1, 'Vladimir': 1, 'indeed,': 1, 'welcome,"': 1,
    'player': 1, 'added': 1, 'Games,': 1, 'Russi...': 1
})

# Option 1: Get a new Counter with only words that appear >1 times
filtered_counter = Counter({word: count for word, count in word_counts.items() if count > 1})

# Option 2: Get just the list of words that appear >1 times (no counts)
filtered_words = [word for word, count in word_counts.items() if count > 1]

# Print the results
print("Filtered Counter (words with count >1):")
print(filtered_counter)
print("\nList of words that appear more than once:")
print(filtered_words)

What this does:

  • Option 1 uses a dictionary comprehension to loop through each item in your original Counter, keeping only those where the count is greater than 1. Wrapping the result back in Counter() lets you keep using all the handy built-in Counter methods.
  • Option 2 gives you a simple list of qualifying words, perfect if you don't need to retain their frequency counts.

Bonus: Fixing Punctuation Issues (Optional)

I noticed some words have punctuation attached (like tough, or welcome,"). If you want to treat tough and tough, as the same word, add this quick preprocessing step before building your Counter:

import string

def clean_word(word):
    # Strip all punctuation from the word
    cleaned = word.translate(str.maketrans('', '', string.punctuation))
    # Optional: Convert to lowercase to make counts case-insensitive (e.g., "ATM" and "atm" count as one word)
    return cleaned.lower()

# Example usage when processing raw text
raw_text = "Your input document text here..."
words = raw_text.split()
cleaned_words = [clean_word(word) for word in words]
word_counts = Counter(cleaned_words)

This will make your frequency counts more accurate if punctuation doesn't add meaning to your NLP task.

内容的提问来源于stack exchange,提问作者Muhammad Sulaman Toor

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 07:26:18