You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中高效计算各ID的平均随访时长

高效计算每个ID的平均随访时长

步骤1:预处理数据

先将Date列转换为datetime类型,同时按ID和Date去重,避免重复日期干扰计算:

import pandas as pd

# 转换日期类型
table['Date'] = pd.to_datetime(table['Date'])

# 按ID和日期去重,保留唯一的日期记录
table_unique = table.drop_duplicates(subset=['ID', 'Date'])

步骤2:分组计算每个ID的随访间隔

通过分组操作,对每个ID的日期降序排序,计算每个日期与该ID最早日期的间隔天数,最后求出每个ID的间隔平均值:

# 定义分组处理函数
def calculate_followup_stats(group):
    # 对当前ID的日期降序排序
    sorted_dates = group['Date'].sort_values(ascending=False).reset_index(drop=True)
    # 获取该ID的最早日期
    earliest_date = sorted_dates.min()
    # 计算每个日期与最早日期的间隔天数
    followup_days = (sorted_dates - earliest_date).dt.days
    # 返回统计结果
    return pd.Series({
        'ID': group['ID'].iloc[0],
        'followup_intervals': followup_days.tolist(),
        'avg_followup': followup_days.mean()
    })

# 应用分组函数到每个ID
result = table_unique.groupby('ID').apply(calculate_followup_stats).reset_index(drop=True)

步骤3:查看结果

result数据框包含每个ID的所有随访间隔天数列表,以及对应的平均随访时长:

print(result)

# 计算所有ID的整体平均随访时长
overall_avg = result['avg_followup'].mean()
print(f"所有ID的整体平均随访时长:{overall_avg:.2f}天")

方案优势

  • 摒弃逐行迭代的低效循环,利用pandas矢量化分组运算处理数据,大数据集下速度提升明显;
  • 直接调用pandas内置日期函数完成转换与间隔计算,无需手动处理字符串转日期;
  • 提前去重过滤重复数据,减少无效计算量。

内容的提问来源于stack exchange,提问作者Shichimi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.21 05:09:10