You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python高效统计每个句子中目标单词列表的出现次数?

Efficient Word Count in Sentences Using Python

When dealing with large lists of words and sentences, efficiency is non-negotiable. The naive approach of checking each target word against every sentence is painfully slow (O(N*M) time complexity), so we need a smarter, scalable method.

Key Optimization: Use a Set for Instant Lookups

First, convert your target word list into a set—this transforms membership checks from O(M) (slow for lists) to O(1) (near-instant for sets), which is a game-changer when working with thousands of words.

Solution Code

import string

def count_target_words(sentences, target_words):
    # Convert target words to a set for lightning-fast lookups
    word_set = set(target_words)
    # Create a translator to strip all punctuation from words
    punctuation_translator = str.maketrans('', '', string.punctuation)
    
    counts = []
    for sentence in sentences:
        # Split sentence into words, clean each by removing punctuation
        cleaned_words = [word.translate(punctuation_translator) for word in sentence.split()]
        # Count how many cleaned words match our target set
        match_count = sum(1 for word in cleaned_words if word in word_set)
        counts.append(match_count)
    
    return counts

# Example usage
list_of_words = ['car', 'motorcycle', 'tree']
list_of_sentences = ["I have a car, but I don't have a motorcycle", "I like elephants but I don't like lions"]

result = count_target_words(list_of_sentences, list_of_words)
print(result)  # Output: [2, 0]

How It Works

  • Set Conversion: word_set = set(target_words) ensures checking if a word is in our target list takes almost no time, even with thousands of entries.
  • Punctuation Handling: str.maketrans creates a reusable tool to strip punctuation from words (so "car," becomes "car", which correctly matches our target word).
  • Memory-Efficient Counting: Using a generator expression (sum(1 for word ...)) avoids creating unnecessary intermediate lists, keeping memory usage low for large datasets.

Additional Customizations

  • Case Insensitivity: If you want to count matches regardless of case (e.g., "Car" should match "car"), adjust the code to normalize case:
    word_set = set(word.lower() for word in target_words)
    cleaned_words = [word.translate(punctuation_translator).lower() for word in sentence.split()]
    
  • Precise Word Extraction: For handling words with apostrophes (like "don't") without splitting them, use a regex pattern to find whole words:
    import re
    word_pattern = re.compile(r"\b[\w']+\b")
    cleaned_words = word_pattern.findall(sentence)
    
    Just remember to strip leading/trailing apostrophes if needed (e.g., "'hello" becomes "hello").

Why This Is Efficient

This approach runs in O(N*K + M) time, where:

  • N = number of sentences
  • K = average number of words per sentence
  • M = number of target words

This is drastically faster than the naive O(N*M) approach, making it perfect for datasets with thousands of entries.

内容的提问来源于stack exchange,提问作者Aventinus

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 09:47:32