使用Pandas按规则交换DataFrame列并同步关联列数据
Got it, let's solve this problem step by step. You need to reorder your DataFrame rows so that seq1_id always contains genes with 0035 in the segment right after the first underscore, and seq2_id contains those with 0042—plus swap the corresponding length and GC columns whenever you swap the IDs. Here's how to do it:
Problem Recap
Your original DataFrame looks like this:
import pandas as pd data = { 'seq1_id': ['g13600.t1_0042_0042', 'g1464.t1_0035_0042', 'g34744.t1_0042_0035', 'g12096.t1_0035_0042'], 'seq2_id': ['g66097.t1_0035_0035', 'g45594.t1_0042_0035', 'g50055.t1_0035_0035', 'g34020.t1_0042_0035'], 'dN': [0.10, 0.67, 0.08, 0.43], 'length1': [45, 56, 34, 31], 'length2': [67, 34, 76, 54], 'GC1': [0.78, 0.65, 0.75, 0.89], 'GC2': [0.67, 0.87, 0.45, 0.76] } df = pd.DataFrame(data)
And you need to transform it so that:
seq1_idonly has IDs where the segment after the first underscore is0035seq2_idonly has IDs where that segment is0042- When swapping
seq1_idandseq2_id, you also swaplength1↔length2andGC1↔GC2
Solution
We have two approaches: a row-wise function (easy to read) and a vectorized method (faster for large datasets).
Approach 1: Row-wise Function (Readable)
First, define a function that checks each row's ID pattern and swaps values when needed:
def adjust_row(row): # Extract the key segment (0035/0042) from each ID seq1_segment = row['seq1_id'].split('_')[1] seq2_segment = row['seq2_id'].split('_')[1] # Check if we need to swap: seq1 is 0042 and seq2 is 0035 if seq1_segment == '0042' and seq2_segment == '0035': # Swap the ID columns row['seq1_id'], row['seq2_id'] = row['seq2_id'], row['seq1_id'] # Swap the length columns row['length1'], row['length2'] = row['length2'], row['length1'] # Swap the GC columns row['GC1'], row['GC2'] = row['GC2'], row['GC1'] return row
Apply this function to every row in your DataFrame:
df_adjusted = df.apply(adjust_row, axis=1)
Approach 2: Vectorized Method (Efficient for Large Data)
If you're working with a big dataset, use boolean indexing to avoid slow row-wise iteration:
# Create a mask to identify rows that need swapping swap_mask = (df['seq1_id'].str.split('_').str[1] == '0042') & (df['seq2_id'].str.split('_').str[1] == '0035') # Swap the ID columns for matching rows df.loc[swap_mask, ['seq1_id', 'seq2_id']] = df.loc[swap_mask, ['seq2_id', 'seq1_id']].values # Swap the length columns df.loc[swap_mask, ['length1', 'length2']] = df.loc[swap_mask, ['length2', 'length1']].values # Swap the GC columns df.loc[swap_mask, ['GC1', 'GC2']] = df.loc[swap_mask, ['GC2', 'GC1']].values
Result
Both methods will give you the transformed DataFrame aligned with your requirements:
| seq1_id | seq2_id | dN | length1 | length2 | GC1 | GC2 |
|---|---|---|---|---|---|---|
| g66097.t1_0035_0035 | g13600.t1_0042_0042 | 0.10 | 67 | 45 | 0.67 | 0.78 |
| g1464.t1_0035_0042 | g45594.t1_0042_0035 | 0.67 | 56 | 34 | 0.65 | 0.87 |
| g50055.t1_0035_0035 | g34744.t1_0042_0035 | 0.08 | 76 | 34 | 0.45 | 0.75 |
| g12096.t1_0035_0042 | g34020.t1_0042_0035 | 0.43 | 31 | 54 | 0.89 | 0.76 |
(Note: This strictly follows your stated requirement to swap length/GC columns when swapping IDs. Any discrepancy with your provided expected result may be a typo in the example.)
内容的提问来源于stack exchange,提问作者Grendel

