Pandas DataFrame列添加循环问题:DNA序列变异体重复生成的代码修改请求
Fix for Duplicate DNA Variant Generation in Pandas DataFrame
Let's break down what's causing your duplicate variants and fix this step by step:
Key Issues in Your Original Code
- List Reference Instead of Copy: When you do
shitass = someshitorhundred = onehundo, you're just assigning a reference to the original list, not creating a new copy. Modifyingshitasschangessomeshitdirectly, and your "reset" doesn't actually create a fresh clone. - Broken Loop Structure: Your outer
while i < 99loop contains a nestedfor r in range(100)loop that overwrites all 100 columns with the same single variant. This means you only generate one unique variant and copy it everywhere. - Uninitialized Variables: The
jcounter isn't reset to 0 for each new variant, and thehundredlist isn't reinitialized properly, so subsequent variants can't pick from all positions.
Fixed Code
import pandas as pd import random # Original DNA sequence original_sequence = [4, 4, 2, 4, 4, 1, 3, 2, 1, 1, 2, 3, 4, 3, 3, 4, 1, 3, 4, 4, 3, 2, 4, 4, 4, 3, 3, 1, 3, 1, 4, 1, 3, 4, 2, 4, 2, 3, 3, 3, 1, 2, 1, 3, 2, 3, 2, 4, 4, 3, 4, 4, 4, 3, 2, 1, 4, 3, 4, 4, 2, 2, 2, 1, 2, 2, 1, 1, 4, 2, 1, 4, 3, 3, 2, 4, 4, 1, 1, 2, 1, 4, 1, 4, 4, 3, 3, 1, 3, 2, 3, 3, 1, 4, 1, 2, 2, 3, 2, 4] nucleotides = [1, 2, 3, 4] total_variants = 100 sequence_length = len(original_sequence) # Initialize DataFrame with original sequence as first column df = pd.DataFrame({0: original_sequence}) # Generate each variant one by one for variant_num in range(1, total_variants + 1): # Create a fresh copy of the original sequence for this variant variant = original_sequence.copy() # Get random number of mutations (1-5) num_mutations = random.randint(1, 5) # Pick unique random positions to mutate (no repeats in the same variant) mutation_positions = random.sample(range(sequence_length), num_mutations) for pos in mutation_positions: # Get current nucleotide at position, then pick a different one current_nuc = variant[pos] # Filter out the current nucleotide to ensure mutation is different possible_nucs = [n for n in nucleotides if n != current_nuc] variant[pos] = random.choice(possible_nucs) # Add the new variant as a column to the DataFrame df[variant_num] = variant # Verify the result print(df.head())
What Changed & Why
- Proper List Copies: Using
original_sequence.copy()creates a new independent list for each variant, so changes to one variant don't affect others or the original sequence. - Simplified Loop Logic: We loop once per variant (100 times total), generating a unique variant each time and adding it directly to the DataFrame. No nested loops that overwrite columns repeatedly.
- Random Position Selection:
random.sample()picks unique positions for mutations in one go, avoiding the need to manage a shrinkinghundredlist manually. This is cleaner and less error-prone. - Cleaner Mutation Logic: Instead of multiple
if/elifblocks, we filter the nucleotide list to exclude the current value, making the code shorter and easier to maintain.
内容的提问来源于stack exchange,提问作者ronald birb
相关产品推荐
相关产品推荐

