You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从另一DataFrame中选取年龄分布匹配的队列样本数据

问题根因
  • 仅支持严格同年龄匹配,一旦某年龄对应的健康人群数量小于患者组该年龄人数,直接跳过该年龄所有匹配流程,大量可利用的相邻年龄样本被浪费
  • 未实现你需要的±1/±2岁年龄偏差规则,匹配容错性极低,真实数据中年龄分布不均的情况下很容易出现匹配结果极少的问题
  • 年龄遍历顺序为无序的set集合,容易出现大样本量年龄优先占用全部可匹配样本,小样本量年龄无样本可用的资源分配问题
修复实现

首先补全依赖导入,原有示例代码缺失pandas导入:

import pandas as pd
import numpy as np

将匹配逻辑封装为可复用函数,支持自定义年龄偏差、多次抽样:

def age_matching(sick_df, sick_age_col, healthy_df, healthy_age_col, max_age_diff=1):
    """
    年龄匹配函数
    :param sick_df: 患者数据集
    :param sick_age_col: 患者数据集年龄列名
    :param healthy_df: 健康人群数据集
    :param healthy_age_col: 健康人群数据集年龄列名
    :param max_age_diff: 允许的最大年龄偏差,1为±1岁,2为±2岁
    :return: 匹配到的健康人群索引列表
    """
    # 统计患者各年龄需要匹配的样本量
    sick_age_cnt = sick_df[sick_age_col].value_counts().sort_index()
    # 优先处理健康样本少的年龄,提高整体匹配率
    healthy_age_cnt = healthy_df[healthy_age_col].value_counts().to_dict()
    sorted_ages = sorted(sick_age_cnt.index, key=lambda a: healthy_age_cnt.get(a, 0))
    
    used_healthy_idx = set()
    matched_result = []
    
    for age in sorted_ages:
        need_cnt = sick_age_cnt[age]
        # 先找同年龄未使用的样本
        same_age = healthy_df[(healthy_df[healthy_age_col] == age) & (~healthy_df.index.isin(used_healthy_idx))]
        select_cnt = min(need_cnt, len(same_age))
        if select_cnt > 0:
            selected = np.random.choice(same_age.index, select_cnt, replace=False)
            matched_result.extend(selected.tolist())
            used_healthy_idx.update(selected)
            need_cnt -= select_cnt
        # 同年龄不够,从偏差范围内找
        if need_cnt > 0 and max_age_diff > 0:
            # 遍历偏差范围内的年龄,从小到大找
            for diff in range(1, max_age_diff+1):
                for offset in [-diff, diff]:
                    cur_age = age + offset
                    if cur_age < 0:
                        continue
                    diff_age = healthy_df[(healthy_df[healthy_age_col] == cur_age) & (~healthy_df.index.isin(used_healthy_idx))]
                    if len(diff_age) == 0:
                        continue
                    select_diff_cnt = min(need_cnt, len(diff_age))
                    selected_diff = np.random.choice(diff_age.index, select_diff_cnt, replace=False)
                    matched_result.extend(selected_diff.tolist())
                    used_healthy_idx.update(selected_diff)
                    need_cnt -= select_diff_cnt
                    if need_cnt == 0:
                        break
                if need_cnt == 0:
                    break
    return matched_result
使用示例
# 允许±1岁匹配
matched_idx = age_matching(x_df, 'x', y_df, 'y', max_age_diff=1)
# 允许±2岁匹配
matched_idx_2 = age_matching(x_df, 'x', y_df, 'y', max_age_diff=2)

# 多次抽样直接多次调用函数即可
for i in range(10):
    cur_matched = age_matching(x_df, 'x', y_df, 'y', max_age_diff=1)
    # 后续处理当前次匹配的样本
效果验证

匹配完成后可通过以下代码对比年龄分布一致性:

# 提取匹配后的健康人群年龄
matched_age = y_df.loc[matched_idx, 'y']
# 打印患者和匹配后健康人群的年龄分布统计
print("患者年龄分布:\n", x_df['x'].describe())
print("匹配后健康人群年龄分布:\n", matched_age.describe())

内容的提问来源于stack exchange,提问作者huanpops

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.26 08:36:03