如何用Python高效统计每个句子中目标单词列表的出现次数?
Efficient Word Count in Sentences Using Python
When dealing with large lists of words and sentences, efficiency is non-negotiable. The naive approach of checking each target word against every sentence is painfully slow (O(N*M) time complexity), so we need a smarter, scalable method.
Key Optimization: Use a Set for Instant Lookups
First, convert your target word list into a set—this transforms membership checks from O(M) (slow for lists) to O(1) (near-instant for sets), which is a game-changer when working with thousands of words.
Solution Code
import string def count_target_words(sentences, target_words): # Convert target words to a set for lightning-fast lookups word_set = set(target_words) # Create a translator to strip all punctuation from words punctuation_translator = str.maketrans('', '', string.punctuation) counts = [] for sentence in sentences: # Split sentence into words, clean each by removing punctuation cleaned_words = [word.translate(punctuation_translator) for word in sentence.split()] # Count how many cleaned words match our target set match_count = sum(1 for word in cleaned_words if word in word_set) counts.append(match_count) return counts # Example usage list_of_words = ['car', 'motorcycle', 'tree'] list_of_sentences = ["I have a car, but I don't have a motorcycle", "I like elephants but I don't like lions"] result = count_target_words(list_of_sentences, list_of_words) print(result) # Output: [2, 0]
How It Works
- Set Conversion:
word_set = set(target_words)ensures checking if a word is in our target list takes almost no time, even with thousands of entries. - Punctuation Handling:
str.maketranscreates a reusable tool to strip punctuation from words (so "car," becomes "car", which correctly matches our target word). - Memory-Efficient Counting: Using a generator expression (
sum(1 for word ...)) avoids creating unnecessary intermediate lists, keeping memory usage low for large datasets.
Additional Customizations
- Case Insensitivity: If you want to count matches regardless of case (e.g., "Car" should match "car"), adjust the code to normalize case:
word_set = set(word.lower() for word in target_words) cleaned_words = [word.translate(punctuation_translator).lower() for word in sentence.split()] - Precise Word Extraction: For handling words with apostrophes (like "don't") without splitting them, use a regex pattern to find whole words:
Just remember to strip leading/trailing apostrophes if needed (e.g., "'hello" becomes "hello").import re word_pattern = re.compile(r"\b[\w']+\b") cleaned_words = word_pattern.findall(sentence)
Why This Is Efficient
This approach runs in O(N*K + M) time, where:
- N = number of sentences
- K = average number of words per sentence
- M = number of target words
This is drastically faster than the naive O(N*M) approach, making it perfect for datasets with thousands of entries.
内容的提问来源于stack exchange,提问作者Aventinus
相关产品推荐
相关产品推荐

