You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

从带填充的DNA序列中提取子序列的高效实现问询

问题描述

我有一个包含padded_seq和seq两列的DataFrame,其中padded_seq为带填充的DNA序列,seq为目标序列。示例DataFrame如下:

padded_seq  seq      
0     AATTGGC  TTG                 
1     AATTGGC  TTG                 
2     AATTGGC  TTG                
3     CGTACGC  TAC                 
4     CGTACGC  TAC                 
5     CGTACGC  TAC                

我需要基于seq列,提取padded_seq中每个字符周围4个核苷酸的窗口序列(即±2个核苷酸),为seq的每个字符生成新行。预期输出如下:

padded_seq  seq      single_letter_padded (separated for clarity)
0     AATTGGC  TTG                 AA T TG
0     AATTGGC  TTG                 AT T GG
0     AATTGGC  TTG                 TT G GC
1     CGTACGC  TAC                 CG T AC
1     CGTACGC  TAC                 GT A CG
1     CGTACGC  TAC                 TA C GC

由于数据量达数百万行,当前代码运行时内核崩溃,代码如下:

# Function to extract sequences based on the given window size
half_window=250 #250 per side

def extract_sequences(df, half_window=250):
    padded_single_nucleotides_list = []
    for y in range(len(df['padded_sequences'])):
        for x in range(df['seq_len'][y]):
            padded_single_nucleotides_list.append(df['padded_sequences'][y][(half_window+x)-half_window:(half_window+x)+half_window+1])
            
    return padded_single_nucleotides_list

#to make sure the seq_df is the correct num of rows-- shouldn't have to do this (use explode?), not sure how to get around it
seq_df=seq_df.loc[np.repeat(seq_df.index, seq_df['seq_len'])].reset_index(drop=True)

# Apply the function to each row and explode the list into separate rows
seq_df['single_nucl_padded'] = extract_sequences(seq_df, half_window=250)

# Display the result
seq_df
高效解决方案

核心思路

摒弃嵌套循环的低效逻辑,利用pandas的向量化操作+explode展开功能,结合字符串切片批量处理,大幅降低内存占用、提升计算效率。

代码实现

import pandas as pd

# 示例数据(实际替换为你的数据集)
data = {
    'padded_seq': ['AATTGGC']*3 + ['CGTACGC']*3,
    'seq': ['TTG']*3 + ['TAC']*3
}
seq_df = pd.DataFrame(data)

half_window = 2  # 示例用±2,实际可改为250

def get_window_sequences(row):
    padded_seq = row['padded_seq']
    target_len = len(row['seq'])
    # 按原代码逻辑,目标seq在padded_seq中的起始位置为half_window
    start_pos = half_window
    # 生成每个目标字符的中心位置
    positions = range(start_pos, start_pos + target_len)
    # 批量提取窗口序列
    return [padded_seq[pos - half_window : pos + half_window + 1] for pos in positions]

# 为每行生成窗口序列列表
seq_df['single_nucl_padded'] = seq_df.apply(get_window_sequences, axis=1)
# 展开列表为单独行
seq_df = seq_df.explode('single_nucl_padded').reset_index(drop=True)

# 可选:生成分隔后的可视化列(仅用于验证结果)
seq_df['single_letter_padded'] = seq_df['single_nucl_padded'].apply(lambda s: f"{s[:half_window]} {s[half_window]} {s[half_window+1:]}")

print(seq_df)

超大数据量优化(百万+行)

如果数据集大到内存无法一次性承载,改用分块处理:

chunk_size = 10000  # 根据内存容量调整大小
output_df = pd.DataFrame()

# 分块读取并处理
for chunk in pd.read_csv('your_raw_data.csv', chunksize=chunk_size):
    chunk['single_nucl_padded'] = chunk.apply(get_window_sequences, axis=1)
    chunk = chunk.explode('single_nucl_padded')
    output_df = pd.concat([output_df, chunk], ignore_index=True)

# 保存结果
output_df.to_csv('processed_result.csv', index=False)

优化说明

  1. 避免了原代码中嵌套循环遍历所有行和字符的低效操作,改用apply对每行批量生成窗口序列,再通过explode展开,内存利用率提升显著。
  2. 原代码中创建超大列表会导致内存峰值过高,这里每行生成小列表后直接展开,减少内存压力。
  3. 分块处理方案可应对千万级以上数据,避免一次性加载全部数据到内存。

内容的提问来源于stack exchange,提问作者youtube

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.05 01:17:36