Python中高效判断字符串不在字符串列表的优化方法
Great question! The slowdown you're seeing comes down to how Python handles membership checks on lists: when you use string not in list_of_sentences, it does a linear search—scanning every element in the list one by one until it either finds a match or reaches the end. For small lists this is fine, but as your list grows (especially with thousands or millions of elements), this gets exponentially slower (it has a time complexity of O(n)).
The Fix: Use a set for Fast Lookups
Python's set data structure is built for exactly this kind of problem. It uses a hash table under the hood, which means membership checks (in/not in) and insertions are average O(1) time—they take roughly the same amount of work no matter how big the set gets.
Here's how to adjust your code to use a set alongside your list (if you need to keep the original order of elements):
- Initialize a set from your existing list to track seen sentences:
list_of_sentences = ["How are you", "Who are you", "Are you okay"] seen_sentences = set(list_of_sentences)
- Update both the list and set when adding new strings:
string = "What about you?" if string not in seen_sentences: list_of_sentences.append(string) seen_sentences.add(string) # Critical: keep the set in sync with the list
Bonus: If You Don't Need to Preserve Order
If the order of sentences doesn't matter, you can skip the list entirely and just use a set. Adding elements becomes even simpler, and duplicates are automatically ignored:
sentence_set = {"How are you", "Who are you", "Are you okay"} string = "What about you?" sentence_set.add(string) # Does nothing if the string is already present
Key Notes
- Strings are hashable, so they work perfectly in sets. If you were working with unhashable types (like lists), you'd need to convert them to something hashable first (e.g., tuples).
- If your initial
list_of_sentenceshas duplicate entries, converting it to a set will automatically remove them. If you want to keep duplicates in the list but only prevent adding new duplicates, the list+set approach still works—just make sure you initialize the set with the unique elements from the list.
内容的提问来源于stack exchange,提问作者Narendra

