如何实现从包含子列表的列表中提取重复元素的通用函数?
Finding Duplicate Elements Across Nested Lists
Got it, let's solve this problem where we need to extract elements that appear both in the top-level of a list and inside any of its nested sublists. The goal is to return these unique elements, preserving their order of first occurrence in the main list.
Here's a Python function that does exactly that:
def find_duplicates_in_nested_lists(input_list): # Split input into main elements and collect all sublist elements main_elements = [] sub_elements = set() for item in input_list: if isinstance(item, list): # Add all elements from sublists to a set (automatically removes duplicates) sub_elements.update(item) else: # Collect top-level non-list elements main_elements.append(item) # Now find elements that exist in both, keeping unique and order of first occurrence duplicates = [] seen = set() for elem in main_elements: if elem in sub_elements and elem not in seen: duplicates.append(elem) seen.add(elem) return duplicates
Let's test it with your examples:
Example 1:
country_list = ['pakistan', 'india', 'iran', 'china', 'afghanistan', ['pakistan', 'iran', 'india']] print(find_duplicates_in_nested_lists(country_list)) # Output: ['pakistan', 'india', 'iran']
Example 2:
countries = ['pakistan', 'india', 'iran', 'china', 'afghanistan', ['pakistan', 'india'], ['china', 'pakistan']] print(find_duplicates_in_nested_lists(countries)) # Output: ['pakistan', 'india', 'china']
How it works:
- Separate elements: We first loop through the input list. For each item, if it's a sublist, we add all its elements to a set (this eliminates duplicates within sublists). If it's a regular element, we add it to a list of top-level elements.
- Find overlaps: We then iterate through the top-level elements, checking if each one exists in the sublist elements set. We use a
seenset to ensure we only add each duplicate element once, even if it appears multiple times in the main list. - Preserve order: By iterating through the main elements list in order, we keep the same sequence as the first occurrence of each duplicate element in the original list.
Notes:
- If you don't care about the order of the output, you could simplify the final step to
return list(set(main_elements) & sub_elements), but this would give an unordered list. - This function handles any number of top-level sublists as shown in your second example. If you need to support deeply nested sublists (like lists inside lists inside lists), we'd need a recursive tweak—just let me know if that's required!
内容的提问来源于stack exchange,提问作者Azeem
相关产品推荐
相关产品推荐

