You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python实现迭代并动态更新DataFrame的R复杂算法?

问题描述

我正尝试将R中的一个复杂算法用Python重写。该算法操作两个结构一致的DataFrame(short和long),二者仅size列范围不同(short为10K-3M,long大于3M),且行数不同。算法会迭代对比short与long,并原地修改long,使每次计算都基于更新后的long。

原始数据来源的DataFramebins结构如下:

chrom chr_arm start end ratio ratio_median size
1     p       1     5   0.3   0.7          5
1     p       1     5   0.9   0.7          5

算法核心步骤

  • 遍历short的每一行
  • 对short的每一行,遍历long的每一行
  • 坐标对比:若short完全在long上游(short.end < long.start),则检查通过
  • 检查通过后:
    1. 计算short与long的ratio_median差值的绝对值
    2. 若差值小于threshold:
      • 用short的start和long的end构造新行
      • 计算新行与bins的重叠区域,取重叠区域的ratio中位数作为新的ratio_median
      • 用新行替换long当前行
    3. 若差值大于threshold:
      • 将short当前行插入到long当前行之后
  • 直至遍历完short所有行

当前Python实现的问题

原R代码依赖嵌套循环和原地修改,我尝试的Python版本如下:

def adjust_segments(short: pd.DataFrame, long: pd.DataFrame) -> pd.DataFrame:
    dest_long = long.copy()
    for gid, group in short.groupby(["chrom", "chr_arm"], sort=False):  # 减少循环次数
        for rowid, row in group.iterrows():
            for long_id, long_row in long.iterrows():
                if (row.end < long_row.start):  # 此处为简化检查,实际有多项
                    if abs(row.ratio_median - long_row.ratio_median) < threshold:
                        new_line = recompute_ratio(row, long_row, bins) 
                        dest_long.loc[long_id, :] = new_line
                    else:
                         # 省略了是否在末尾等检查
                        dest_long.loc[long_id + 1, :] = new_line 

该实现无法在每次迭代时使用更新后的long数据,且Pandas中频繁原地修改/插入行效率极低,需要更Pythonic的方案。


Pythonic实现方案

1. 用列表维护long行数据,规避DataFrame低效修改

Pandas是列存储结构,插入、修改行的性能远低于原生列表。可以先将long转为字典列表,遍历过程中直接操作列表,最后再转回DataFrame:

import pandas as pd

def adjust_segments(short: pd.DataFrame, long: pd.DataFrame, bins: pd.DataFrame, threshold: float) -> pd.DataFrame:
    # 将long转为字典列表,支持动态插入/修改
    long_rows = long.to_dict('records')
    # 按染色体+臂分组,减少跨区域无效对比
    short_groups = short.groupby(["chrom", "chr_arm"], sort=False)
    # 预分组bins,提升重叠区域查询效率
    bins_groups = bins.groupby(["chrom", "chr_arm"])

    for (chrom, arm), short_group in short_groups:
        # 过滤当前染色体臂对应的long行索引
        filtered_long_idx = [idx for idx, row in enumerate(long_rows) 
                             if row['chrom'] == chrom and row['chr_arm'] == arm]
        # 遍历当前组的short行
        for _, short_row in short_group.iterrows():
            # 倒序遍历long行,避免插入新行后打乱未遍历的索引
            for long_idx in reversed(filtered_long_idx):
                long_row = long_rows[long_idx]
                # 坐标检查(可扩展为实际的多项校验逻辑)
                if short_row['end'] < long_row['start']:
                    diff = abs(short_row['ratio_median'] - long_row['ratio_median'])
                    if diff < threshold:
                        # 构造新行并替换原long行
                        new_start = short_row['start']
                        new_end = long_row['end']
                        # 计算重叠区域的ratio中位数
                        current_bins = bins_groups.get_group((chrom, arm))
                        overlap = current_bins[
                            (current_bins['start'] < new_end) &
                            (current_bins['end'] > new_start)
                        ]
                        new_ratio_median = overlap['ratio'].median()
                        # 更新新行字段
                        new_row = long_row.copy()
                        new_row.update({
                            'start': new_start,
                            'end': new_end,
                            'ratio_median': new_ratio_median,
                            'size': new_end - new_start
                        })
                        long_rows[long_idx] = new_row
                    else:
                        # 插入short行到当前long行之后
                        long_rows.insert(long_idx + 1, short_row.to_dict())
                        # 更新过滤后的long索引,适配插入操作
                        filtered_long_idx = [idx if idx <= long_idx else idx + 1 
                                            for idx in filtered_long_idx]
                    # 匹配到目标行后跳出循环(按业务逻辑调整)
                    break
    # 转回DataFrame返回
    return pd.DataFrame(long_rows)

2. 关键优化点说明

  • 倒序遍历long行:插入新行时不会干扰未遍历的行索引,省去频繁维护索引的麻烦
  • 预分组过滤:对short、long、bins都按chrom+chr_arm分组,大幅减少无效遍历和查询
  • 列表操作优先:利用Python列表的O(1)修改和O(n)插入(局部场景下性能远超DataFrame)

内容的提问来源于stack exchange,提问作者Einar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.16 00:54:59