You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Pandas按规则交换DataFrame列并同步关联列数据

Pandas: Adjust Rows to Enforce ID Patterns and Swap Corresponding Columns

Got it, let's solve this problem step by step. You need to reorder your DataFrame rows so that seq1_id always contains genes with 0035 in the segment right after the first underscore, and seq2_id contains those with 0042—plus swap the corresponding length and GC columns whenever you swap the IDs. Here's how to do it:

Problem Recap

Your original DataFrame looks like this:

import pandas as pd

data = {
    'seq1_id': ['g13600.t1_0042_0042', 'g1464.t1_0035_0042', 'g34744.t1_0042_0035', 'g12096.t1_0035_0042'],
    'seq2_id': ['g66097.t1_0035_0035', 'g45594.t1_0042_0035', 'g50055.t1_0035_0035', 'g34020.t1_0042_0035'],
    'dN': [0.10, 0.67, 0.08, 0.43],
    'length1': [45, 56, 34, 31],
    'length2': [67, 34, 76, 54],
    'GC1': [0.78, 0.65, 0.75, 0.89],
    'GC2': [0.67, 0.87, 0.45, 0.76]
}
df = pd.DataFrame(data)

And you need to transform it so that:

  • seq1_id only has IDs where the segment after the first underscore is 0035
  • seq2_id only has IDs where that segment is 0042
  • When swapping seq1_id and seq2_id, you also swap length1 ↔ length2 and GC1 ↔ GC2

Solution

We have two approaches: a row-wise function (easy to read) and a vectorized method (faster for large datasets).

Approach 1: Row-wise Function (Readable)

First, define a function that checks each row's ID pattern and swaps values when needed:

def adjust_row(row):
    # Extract the key segment (0035/0042) from each ID
    seq1_segment = row['seq1_id'].split('_')[1]
    seq2_segment = row['seq2_id'].split('_')[1]
    
    # Check if we need to swap: seq1 is 0042 and seq2 is 0035
    if seq1_segment == '0042' and seq2_segment == '0035':
        # Swap the ID columns
        row['seq1_id'], row['seq2_id'] = row['seq2_id'], row['seq1_id']
        # Swap the length columns
        row['length1'], row['length2'] = row['length2'], row['length1']
        # Swap the GC columns
        row['GC1'], row['GC2'] = row['GC2'], row['GC1']
    return row

Apply this function to every row in your DataFrame:

df_adjusted = df.apply(adjust_row, axis=1)

Approach 2: Vectorized Method (Efficient for Large Data)

If you're working with a big dataset, use boolean indexing to avoid slow row-wise iteration:

# Create a mask to identify rows that need swapping
swap_mask = (df['seq1_id'].str.split('_').str[1] == '0042') & (df['seq2_id'].str.split('_').str[1] == '0035')

# Swap the ID columns for matching rows
df.loc[swap_mask, ['seq1_id', 'seq2_id']] = df.loc[swap_mask, ['seq2_id', 'seq1_id']].values
# Swap the length columns
df.loc[swap_mask, ['length1', 'length2']] = df.loc[swap_mask, ['length2', 'length1']].values
# Swap the GC columns
df.loc[swap_mask, ['GC1', 'GC2']] = df.loc[swap_mask, ['GC2', 'GC1']].values

Result

Both methods will give you the transformed DataFrame aligned with your requirements:

seq1_idseq2_iddNlength1length2GC1GC2
g66097.t1_0035_0035g13600.t1_0042_00420.1067450.670.78
g1464.t1_0035_0042g45594.t1_0042_00350.6756340.650.87
g50055.t1_0035_0035g34744.t1_0042_00350.0876340.450.75
g12096.t1_0035_0042g34020.t1_0042_00350.4331540.890.76

(Note: This strictly follows your stated requirement to swap length/GC columns when swapping IDs. Any discrepancy with your provided expected result may be a typo in the example.)

内容的提问来源于stack exchange,提问作者Grendel

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 07:36:03