You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何实现列表中含特定子串元素的去重?保留首次出现项

How to Deduplicate List Elements Based on Specific Substrings (Keep First Occurrence)

Alright, let's break down how to solve this problem—you want to check if specific substrings from a list appear in multiple elements of another list, then keep only the first element that contains each substring and remove all subsequent duplicates. Here's a straightforward, efficient way to do this in Python:

Step-by-Step Approach

  1. Track seen substrings: Use a set to keep track of which target substrings we've already encountered (sets have fast lookup times, which makes this efficient).
  2. Iterate through the original list: For each element, check if it contains any of the target substrings.
  3. Keep or discard elements:
    • If the element contains a substring we haven't seen before: add the element to our result list and mark the substring as seen.
    • If the element contains a substring we've already seen: skip it.
    • If the element has no matching target substrings: keep it as-is (since it doesn't contribute to the duplicate issue).

Code Implementation (Using Your Example)

Let's apply this to your exact sample data:

# Your input data
SUBSTRINGS = ['banana', 'chocolate']
MYLIST = ['1 banana cake', '2 banana cake', '3 cherry cake', '4 chocolate cake', '5 chocolate cake', '6 banana cake', '7 pineapple cake']

# Initialize tracking set and result list
seen_substrings = set()
unique_list = []

for item in MYLIST:
    # Find all target substrings present in the current item
    matched_substrings = [sub for sub in SUBSTRINGS if sub in item]
    
    if matched_substrings:
        # Grab the first matching substring (adjust if you need to handle multiple matches differently)
        first_match = matched_substrings[0]
        if first_match not in seen_substrings:
            seen_substrings.add(first_match)
            unique_list.append(item)
    else:
        # No target substrings found—keep the item
        unique_list.append(item)

# Print the result
print(unique_list)
# Output: ['1 banana cake', '3 cherry cake', '4 chocolate cake', '7 pineapple cake']

Key Notes

  • Handling multiple substrings in one element: If an item contains more than one target substring (e.g., "banana chocolate muffin"), the code above uses the first matching substring from your SUBSTRINGS list. If you need a different behavior (like checking if any of the substrings are unseen), you can adjust the logic to check all matches instead of just the first.
  • Efficiency: Using a set for seen_substrings ensures that checking if a substring is already seen is an O(1) operation, making the overall process O(n*m) where n is the length of MYLIST and m is the length of SUBSTRINGS—this is very efficient for most use cases.

内容的提问来源于stack exchange,提问作者douglas780

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 06:52:08