如何优化处理相同翻译源对应不同结果的Python列表代码?
Hey there! I see you're looking to streamline your code that tracks duplicate source entries and their corresponding differing target translations. The nested loops in your current implementation can get slow with large datasets, so let's refactor this into a cleaner, more efficient solution.
The Core Problem
We have two lists: source (with duplicate entries) and target (where duplicates in source may map to different values). We need to generate a readable structure showing which indexes of a repeated source element link to distinct target values.
Optimized Implementation
Instead of multiple passes and nested loops, we can use a single traversal to group our data, then process those groups to isolate divergent targets. Here's how:
from collections import defaultdict def map_source_to_divergent_targets(source, target): # First, group all (index, target_value) pairs by their source element source_groups = defaultdict(list) for idx, (src, tgt) in enumerate(zip(source, target)): source_groups[src].append((idx, tgt)) # Now, for each source group, group indexes by their target value result = {} for src, items in source_groups.items(): # Skip sources that only appear once if len(items) <= 1: continue target_index_map = defaultdict(list) for idx, tgt in items: target_index_map[tgt].append(idx) # Only keep sources that have multiple distinct targets if len(target_index_map) > 1: result[src] = target_index_map return result # Example usage source = [1, 1, 2, 3, 4, 4, 4, 5, 6] target = [1, 2, 2, 3, 1, 2, 3, 5, 6] output = map_source_to_divergent_targets(source, target) print(output)
What This Does
- Single Pass Grouping: We iterate through
sourceandtargetonce, pairing each index with its corresponding source/target value and grouping by the source element. This runs in O(n) time, where n is the length of your lists. - Target Index Grouping: For each source that appears multiple times, we then group its indexes by the target value. This lets us instantly see which indexes map to which target.
- Filter for Divergence: We only include sources that have more than one distinct target value in our final result, keeping the output focused on what you care about.
Sample Output
Running the example code above will give you:
{ 1: {1: [0], 2: [1]}, 4: {1: [4], 2: [5], 3: [6]} }
This tells you:
- Source
1maps to target1at index0, and target2at index1 - Source
4maps to target1at index4, target2at index5, and target3at index6
Why This Is Better Than Your Current Code
- Faster: No repeated calls to
list.index()(which is O(n) each time) or nested loops that multiply your time complexity. - Cleaner: The logic is broken into clear, manageable steps that are easy to read and modify.
- More Flexible: The output structure is directly usable (a dictionary of dictionaries) instead of just print statements, making it easier to integrate into other parts of your code.
内容的提问来源于stack exchange,提问作者Hisham

