You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python读取52MB .ped文件内存高、处理慢的优化方案咨询

性能瓶颈核心原因

你当前代码慢的核心原因是每次循环都重新完整读取一遍整个文件:pd.read_csv每次调用都会从磁盘起始位置扫描全量107行内容,跳过非目标列后才加载你需要的2列,重复磁盘IO和重复解析的开销占了总耗时的99%以上。按当前逻辑总共要触发超过12万次全文件扫描,总耗时数小时是必然结果。之前测试dtype='category'反而变慢,是因为单次仅加载2列时,分类编码的构造开销远大于内存优化收益,完全不适用。

注意这个文件总共只有107行,哪怕列数高达24万,全量处理的内存和时间开销都极低,完全不需要逐列读取。

最优实现方案(预计总耗时<10秒)

全程只读1次文件,用Python内置模块边读边计数,不存储冗余数据,内存占用不超过10MB:

import csv
from collections import defaultdict

freq_count = []
VALID_BASES = {'A', 'C', 'G', 'T'}

with open('filtered_conv.ped', 'r', newline='') as f:
    reader = csv.reader(f, delimiter='\t')
    first_row = next(reader)
    total_cols = len(first_row)
    # 初始化每个位点的碱基计数器,从第6列(索引6)开始每两列对应一个位点
    site_counters = [defaultdict(int) for _ in range(6, total_cols, 2)]
    site_positions = list(range(6, total_cols, 2))

    # 处理第一行数据
    for site_idx, col_start in enumerate(site_positions):
        b1, b2 = first_row[col_start], first_row[col_start+1]
        if b1 in VALID_BASES:
            site_counters[site_idx][b1] += 1
        if b2 in VALID_BASES:
            site_counters[site_idx][b2] += 1

    # 逐行处理剩余样本
    for row in reader:
        for site_idx, col_start in enumerate(site_positions):
            b1, b2 = row[col_start], row[col_start+1]
            if b1 in VALID_BASES:
                site_counters[site_idx][b1] += 1
            if b2 in VALID_BASES:
                site_counters[site_idx][b2] += 1

# 整理为你需要的结果格式
for counter in site_counters:
    # 按出现次数降序,最高的为ref,次高的为alt
    sorted_bases = sorted(counter.items(), key=lambda x: -x[1])
    ref_nb, c_ref = sorted_bases[0]
    if len(sorted_bases) >= 2:
        alt_nb, c_alt = sorted_bases[1]
    else:
        alt_nb, c_alt = None, 0
    freq_count.append({
        'ref_Nb': ref_nb,
        'c_ref': c_ref,
        'alt_Nb': alt_nb,
        'c_alt': c_alt
    })

方案优势

  • 仅触发1次磁盘读取,无重复IO开销
  • 无pandas DataFrame构造、索引转换等额外固定开销,计数逻辑直接在读取时完成
  • 内存中仅存储每个位点的碱基计数值,不需要缓存全量基因型数据
基于pandas的快速方案(预计总耗时<30秒)

如果习惯pandas生态,一次性加载所有需要的列,不要逐列读取,利用category类型压缩内存:

import pandas as pd

# 一次性加载所有需要的列,跳过前6列,用category类型压缩内存
df = pd.read_csv(
    'filtered_conv.ped',
    sep='\t',
    header=None,
    dtype='category',
    usecols=range(6, 248189)
)

# 拆分两个等位基因对应的列
allele1_cols = df.columns[::2]
allele2_cols = df.columns[1::2]

# 批量统计所有列的碱基计数
count1 = df[allele1_cols].apply(pd.value_counts).fillna(0)
count2 = df[allele2_cols].apply(pd.value_counts).fillna(0)
total_count = count1.add(count2, fill_value=0)

# 整理结果格式
freq_count = []
for col in total_count.columns:
    col_cnt = total_count[col].sort_values(ascending=False)
    ref_nb = col_cnt.index[0]
    c_ref = int(col_cnt.iloc[0])
    if len(col_cnt) >= 2:
        alt_nb = col_cnt.index[1]
        c_alt = int(col_cnt.iloc[1])
    else:
        alt_nb, c_alt = None, 0
    freq_count.append({
        'ref_Nb': ref_nb,
        'c_ref': c_ref,
        'alt_Nb': alt_nb,
        'c_alt': c_alt
    })

这个方案内存占用约100MB,批量操作的开销远低于逐列循环读取。


内容的提问来源于stack exchange,提问作者Igor Salerno Filgueiras

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.26 13:33:24