如何生成无重复及对称重复的itertools product结果?
itertools.product I get exactly what you're trying to do here—handling product results across multiple lists while avoiding duplicate values and redundant symmetric pairs, all without killing performance. Let's break down a solution that meets all your requirements:
Core Approach
We'll use a generator to process items on-the-fly (so we don't load the entire product into memory at once) plus a set to track unique "signature keys" for each valid item. Here's the breakdown:
- Filter out items with duplicate values: Skip any tuple where an element appears more than once.
- Deduplicate symmetric pairs: For each valid item, create a unique key that preserves the first element (since your symmetric example keeps the
add/suboperator fixed) but sorts the remaining elements. This way, pairs like('add',3,'a')and('add','a',3)will generate the same key, and we only keep the first occurrence.
Implementation Code
import itertools def has_duplicates(seq): """Quickly check if a sequence has duplicate elements (stops early on hit)""" seen = set() for elem in seq: if elem in seen: return True seen.add(elem) return False def filtered_product(*args): seen_keys = set() for item in itertools.product(*args): # Skip items with duplicate values if has_duplicates(item): continue # Create a unique key: first element + sorted rest of the elements signature = (item[0],) + tuple(sorted(item[1:])) if signature not in seen_keys: seen_keys.add(signature) yield item
Test It With Your Example
a = ['add', 'sub'] b = [2,3, 'a', 'b'] c = [1,3, 'a', 'c'] # Get the filtered result result = list(filtered_product(a, b, c))
Why This Works for Your Requirements
- No duplicate values: The
has_duplicatesfunction stops as soon as it finds a repeat, making it faster than creating a full set for every tuple. - No symmetric pairs: The sorted signature ensures that any permutation of the non-first elements gets mapped to the same key, so we only keep the first occurrence.
- Multi-list compatibility: Unlike
combinationswhich works on a single list, this handles any number of input lists (5+ is no problem). - Performance-friendly: Using a generator means we never store the entire product in memory, and set lookups are O(1), so the overhead stays low even with large input lists.
Optional Optimization
If your input lists have a lot of duplicate elements upfront, you could pre-process each list to deduplicate first (e.g., a = list(set(a))) to reduce the total number of product items we need to process in the first place.
内容的提问来源于stack exchange,提问作者xyhuang

