如何基于重叠染色体坐标将字典拆分为字典列表?
Solution: Group Dictionary Keys by Overlapping Chromosome Intervals
Let's break down how to solve this problem effectively. The goal is to group keys from your dictionary into clusters where their coordinate intervals (using min/max values) overlap by at least 1000 base pairs, while keeping separate chromosomes distinct. Here's a step-by-step solution with Python code:
Step 1: Understand the Core Requirements
We need to:
- Group by chromosome: Never mix entries from different chromosomes.
- Calculate interval bounds: For each key, use the minimum and maximum coordinates from its value list.
- Merge overlapping intervals: Combine keys if their intervals share at least 1000 base pairs; keep non-overlapping keys in their own dictionaries.
- Assemble into a list: Convert each merged group (and standalone keys) into a dictionary, then collect all into
New_list.
Step 2: Python Implementation
from collections import defaultdict # Your original dictionary First_dict = { 'Key1': ['chr10', 19010495, 19014590, 19014064], 'Key2': ['chr10', 19010495, 19014658], 'Key3': ['chr10', 19010502, 19014641], 'Key4': ['chr10', 37375766, 37377526], 'Key5': ['chr10', 76310389, 76315990, 76312224, 76312963], 'Key6': ['chr11', 14806147, 14814006] } # 1. Group entries by chromosome, precomputing min/max coordinates chr_groups = defaultdict(list) for key, coords in First_dict.items(): chr_name = coords[0] min_coord = min(coords[1:]) max_coord = max(coords[1:]) chr_groups[chr_name].append((key, min_coord, max_coord)) # 2. Process each chromosome group to merge overlapping intervals New_list = [] for chr_name, entries in chr_groups.items(): # Sort entries by their starting coordinate to simplify merging sorted_entries = sorted(entries, key=lambda x: x[1]) merged_groups = [] if sorted_entries: # Initialize with the first entry current_group = [sorted_entries[0][0]] current_min = sorted_entries[0][1] current_max = sorted_entries[0][2] for key, min_coord, max_coord in sorted_entries[1:]: # Calculate overlap length between current group and new entry overlap_start = max(current_min, min_coord) overlap_end = min(current_max, max_coord) overlap_length = overlap_end - overlap_start + 1 if overlap_end >= overlap_start else 0 if overlap_length >= 1000: # Merge into current group and update interval bounds current_group.append(key) current_min = min(current_min, min_coord) current_max = max(current_max, max_coord) else: # Finalize current group and start a new one merged_groups.append(current_group) current_group = [key] current_min = min_coord current_max = max_coord # Add the last remaining group merged_groups.append(current_group) # 3. Convert merged groups to dictionaries and add to New_list for group in merged_groups: New_list.append({k: First_dict[k] for k in group}) # Optional: Print the result to verify for idx, group_dict in enumerate(New_list, 1): print(f"Dictionary {idx}:") for key, coords in group_dict.items(): print(f" {key}: {coords}") print()
Step 3: Expected Output
Running this code will produce New_list with 4 dictionaries:
- A dictionary containing
Key1,Key2,Key3(their chr10 intervals overlap by far more than 1000 bp) - A dictionary with just
Key4(chr10 interval doesn't overlap with others) - A dictionary with just
Key5(chr10 interval is isolated) - A dictionary with just
Key6(chr11 interval is standalone)
Key Details Explained
- Chromosome grouping: We use
defaultdictto quickly organize entries by their chromosome name, ensuring no cross-chromosome merging. - Interval sorting: Sorting entries by start coordinate makes merging straightforward—we only need to compare each entry to the current merged group.
- Overlap calculation: We explicitly compute the length of the overlap between intervals to ensure it meets the 1000 bp threshold.
内容的提问来源于stack exchange,提问作者RAHenriksen
相关产品推荐
相关产品推荐

