Python读取52MB .ped文件内存高、处理慢的优化方案咨询
性能瓶颈核心原因
你当前代码慢的核心原因是每次循环都重新完整读取一遍整个文件:pd.read_csv每次调用都会从磁盘起始位置扫描全量107行内容,跳过非目标列后才加载你需要的2列,重复磁盘IO和重复解析的开销占了总耗时的99%以上。按当前逻辑总共要触发超过12万次全文件扫描,总耗时数小时是必然结果。之前测试dtype='category'反而变慢,是因为单次仅加载2列时,分类编码的构造开销远大于内存优化收益,完全不适用。
注意这个文件总共只有107行,哪怕列数高达24万,全量处理的内存和时间开销都极低,完全不需要逐列读取。
最优实现方案(预计总耗时<10秒)
全程只读1次文件,用Python内置模块边读边计数,不存储冗余数据,内存占用不超过10MB:
import csv from collections import defaultdict freq_count = [] VALID_BASES = {'A', 'C', 'G', 'T'} with open('filtered_conv.ped', 'r', newline='') as f: reader = csv.reader(f, delimiter='\t') first_row = next(reader) total_cols = len(first_row) # 初始化每个位点的碱基计数器,从第6列(索引6)开始每两列对应一个位点 site_counters = [defaultdict(int) for _ in range(6, total_cols, 2)] site_positions = list(range(6, total_cols, 2)) # 处理第一行数据 for site_idx, col_start in enumerate(site_positions): b1, b2 = first_row[col_start], first_row[col_start+1] if b1 in VALID_BASES: site_counters[site_idx][b1] += 1 if b2 in VALID_BASES: site_counters[site_idx][b2] += 1 # 逐行处理剩余样本 for row in reader: for site_idx, col_start in enumerate(site_positions): b1, b2 = row[col_start], row[col_start+1] if b1 in VALID_BASES: site_counters[site_idx][b1] += 1 if b2 in VALID_BASES: site_counters[site_idx][b2] += 1 # 整理为你需要的结果格式 for counter in site_counters: # 按出现次数降序,最高的为ref,次高的为alt sorted_bases = sorted(counter.items(), key=lambda x: -x[1]) ref_nb, c_ref = sorted_bases[0] if len(sorted_bases) >= 2: alt_nb, c_alt = sorted_bases[1] else: alt_nb, c_alt = None, 0 freq_count.append({ 'ref_Nb': ref_nb, 'c_ref': c_ref, 'alt_Nb': alt_nb, 'c_alt': c_alt })
方案优势
- 仅触发1次磁盘读取,无重复IO开销
- 无pandas DataFrame构造、索引转换等额外固定开销,计数逻辑直接在读取时完成
- 内存中仅存储每个位点的碱基计数值,不需要缓存全量基因型数据
基于pandas的快速方案(预计总耗时<30秒)
如果习惯pandas生态,一次性加载所有需要的列,不要逐列读取,利用category类型压缩内存:
import pandas as pd # 一次性加载所有需要的列,跳过前6列,用category类型压缩内存 df = pd.read_csv( 'filtered_conv.ped', sep='\t', header=None, dtype='category', usecols=range(6, 248189) ) # 拆分两个等位基因对应的列 allele1_cols = df.columns[::2] allele2_cols = df.columns[1::2] # 批量统计所有列的碱基计数 count1 = df[allele1_cols].apply(pd.value_counts).fillna(0) count2 = df[allele2_cols].apply(pd.value_counts).fillna(0) total_count = count1.add(count2, fill_value=0) # 整理结果格式 freq_count = [] for col in total_count.columns: col_cnt = total_count[col].sort_values(ascending=False) ref_nb = col_cnt.index[0] c_ref = int(col_cnt.iloc[0]) if len(col_cnt) >= 2: alt_nb = col_cnt.index[1] c_alt = int(col_cnt.iloc[1]) else: alt_nb, c_alt = None, 0 freq_count.append({ 'ref_Nb': ref_nb, 'c_ref': c_ref, 'alt_Nb': alt_nb, 'c_alt': c_alt })
这个方案内存占用约100MB,批量操作的开销远低于逐列循环读取。
内容的提问来源于stack exchange,提问作者Igor Salerno Filgueiras
相关产品推荐
相关产品推荐

