You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python如何对pandas大样本下采样以匹配另一数据集的变量分布

匹配身高分布的抽样实现方案

pandas没有直接实现该需求的内置函数,但可以通过以下两种常用方案快速实现:

方案1:分位数分桶匹配抽样

该方案通过匹配身高区间的样本占比实现分布对齐,匹配精度高,可灵活调整分桶粒度控制匹配效果:

import pandas as pd
import numpy as np

# 1. 基于df2的身高生成分桶边界,示例用5分位,可调整为10分位提升匹配精度
q_bins = df_2['height'].quantile([0, 0.2, 0.4, 0.6, 0.8, 1]).values
# 2. 统计df2每个身高桶的样本量,即需要从df1对应桶抽取的数量
df_2['height_bin'] = pd.cut(df_2['height'], bins=q_bins, include_lowest=True)
bin_need_count = df_2['height_bin'].value_counts().sort_index()

# 3. 给df1打相同的身高桶标签
df_1['height_bin'] = pd.cut(df_1['height'], bins=q_bins, include_lowest=True)

# 4. 按桶抽样后合并
sampled_bins = []
for bin_range, need_num in bin_need_count.items():
    bin_df = df_1[df_1['height_bin'] == bin_range]
    # 若对应桶样本不足可降低分桶粒度,或临时开启replace=True放回抽样
    sampled_bin = bin_df.sample(n=need_num, random_state=42, replace=False)
    sampled_bins.append(sampled_bin)

# 得到最终匹配分布的抽样子集
df_sampled = pd.concat(sampled_bins).drop(columns='height_bin').reset_index(drop=True)

方案2:加权概率抽样

如果只需要匹配df2的均值、标准差等整体统计特征,可直接用正态分布概率加权抽样,实现更简单:

import pandas as pd
from scipy.stats import norm

# 1. 读取df2的身高统计量
mu, std = df_2['height'].mean(), df_2['height'].std()
# 2. 给df1每个样本计算抽样权重:越接近df2身高分布的样本权重越高
df_1['weight'] = norm.pdf(df_1['height'], loc=mu, scale=std)
# 3. 按权重抽取200条样本
df_sampled = df_1.sample(n=200, weights='weight', random_state=42, replace=False).drop(columns='weight').reset_index(drop=True)

匹配效果校验

抽样完成后可直接对比统计量确认匹配效果:

print("df2身高统计量:")
print(df_2['height'].describe()[['mean', 'std', 'min', '50%']])
print("抽样结果身高统计量:")
print(df_sampled['height'].describe()[['mean', 'std', 'min', '50%']])

注意:如果df1的身高范围没有覆盖df2的最小/最大值,会出现匹配度偏差,抽样前建议先校验df1的身高覆盖范围。

内容的提问来源于stack exchange,提问作者tkxgoogle

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.23 17:06:00