如何实现列表中含特定子串元素的去重?保留首次出现项
How to Deduplicate List Elements Based on Specific Substrings (Keep First Occurrence)
Alright, let's break down how to solve this problem—you want to check if specific substrings from a list appear in multiple elements of another list, then keep only the first element that contains each substring and remove all subsequent duplicates. Here's a straightforward, efficient way to do this in Python:
Step-by-Step Approach
- Track seen substrings: Use a set to keep track of which target substrings we've already encountered (sets have fast lookup times, which makes this efficient).
- Iterate through the original list: For each element, check if it contains any of the target substrings.
- Keep or discard elements:
- If the element contains a substring we haven't seen before: add the element to our result list and mark the substring as seen.
- If the element contains a substring we've already seen: skip it.
- If the element has no matching target substrings: keep it as-is (since it doesn't contribute to the duplicate issue).
Code Implementation (Using Your Example)
Let's apply this to your exact sample data:
# Your input data SUBSTRINGS = ['banana', 'chocolate'] MYLIST = ['1 banana cake', '2 banana cake', '3 cherry cake', '4 chocolate cake', '5 chocolate cake', '6 banana cake', '7 pineapple cake'] # Initialize tracking set and result list seen_substrings = set() unique_list = [] for item in MYLIST: # Find all target substrings present in the current item matched_substrings = [sub for sub in SUBSTRINGS if sub in item] if matched_substrings: # Grab the first matching substring (adjust if you need to handle multiple matches differently) first_match = matched_substrings[0] if first_match not in seen_substrings: seen_substrings.add(first_match) unique_list.append(item) else: # No target substrings found—keep the item unique_list.append(item) # Print the result print(unique_list) # Output: ['1 banana cake', '3 cherry cake', '4 chocolate cake', '7 pineapple cake']
Key Notes
- Handling multiple substrings in one element: If an item contains more than one target substring (e.g.,
"banana chocolate muffin"), the code above uses the first matching substring from yourSUBSTRINGSlist. If you need a different behavior (like checking if any of the substrings are unseen), you can adjust the logic to check all matches instead of just the first. - Efficiency: Using a set for
seen_substringsensures that checking if a substring is already seen is an O(1) operation, making the overall process O(n*m) where n is the length ofMYLISTand m is the length ofSUBSTRINGS—this is very efficient for most use cases.
内容的提问来源于stack exchange,提问作者douglas780
相关产品推荐
相关产品推荐

