You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python字典查找同一句子中重复出现的关联词汇?

Find Repeated Words (Associated Vocab) in Each Sentence Using Dictionaries

Hey there! Great job getting the sentences split into word lists already — that's a solid first step. Let's walk through how to use dictionaries to track repeated words per sentence, exactly what you need for your "associated vocab" requirement.

The Core Idea

For each sentence's word list, we'll use a dictionary to count how many times each word appears. Any word with a count greater than 1 is your associated vocabulary for that sentence.

Step-by-Step Implementation

Here's how to build this onto your existing code:

First, let's start with your existing split sentences structure (I'll add a sample sentences list for context, but you can replace it with your actual data):

# Your existing code (with sample sentences for demo)
sentences = [
    "I love Python, Python is fun and I love coding",
    "Hello world, world says hello to everyone",
    "Single words here have no repeats"
]

split_sentences = []
for sentence in sentences:
    # Note: If you want to clean punctuation (like commas), you might add .strip(',.!?') to each word later
    split_sentences.append(sentence.split())

Now, let's add the dictionary-based counting logic:

# Store results: a list where each entry is a dict of {word: count} for repeated words in the sentence
associated_vocab_per_sentence = []

for word_list in split_sentences:
    word_counts = {}
    # Count each word in the current sentence
    for word in word_list:
        # Optional: Use word.lower() if you want case-insensitive counting (e.g., "Python" and "python" count as same)
        # cleaned_word = word.lower().strip(',.!?')
        cleaned_word = word.strip(',.!?')  # Basic punctuation cleanup
        if cleaned_word in word_counts:
            word_counts[cleaned_word] += 1
        else:
            word_counts[cleaned_word] = 1
    
    # Filter to keep only words that appear more than once
    repeated_words = {word: count for word, count in word_counts.items() if count > 1}
    associated_vocab_per_sentence.append(repeated_words)

# Let's print the results to see how it works
for idx, vocab in enumerate(associated_vocab_per_sentence):
    print(f"Sentence {idx+1} Associated Vocab: {vocab}")

What This Does

  • Word Counting: The inner loop goes through each word in the sentence, updating the dictionary to track how many times each word shows up. We added a quick strip(',.!?') to clean up basic punctuation — you can adjust this if your text has other symbols.
  • Filter Repeats: Using a dictionary comprehension, we filter out any words that only appear once, leaving just your associated vocabulary.
  • Results Storage: We collect these results in a list so you can easily access the repeated words for each original sentence.

Sample Output

Running the code above would give you:

Sentence 1 Associated Vocab: {'I': 2, 'love': 2, 'Python': 2}
Sentence 2 Associated Vocab: {'world': 2}  # If using .lower(), this would include {'hello':2, 'world':2}
Sentence 3 Associated Vocab: {}

If you want case-insensitive counting, just swap the cleaned_word line to use word.lower().strip(',.!?') — that would make "Hello" and "hello" count as the same word in sentence 2.

内容的提问来源于stack exchange,提问作者fskoft

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 07:18:38