You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何调整Pandas sample方法对指定列特定类别选择性抽样并保留其他类别数据?

当然可以实现这个需求!用Pandas完全能搞定,而且方法很直观:咱们把数据拆成两部分处理——对需要抽样的类别做抽样,其他类别直接保留全部数据,最后再合并到一起就行。

具体步骤和代码示例

首先,先构造你提供的示例DataFrame:

import pandas as pd

# 构造示例数据
data = {
    'Predicted': [4.561074719, 3.114821134, 5.47200407, 7.048857494, 5.318448093,
                  3.81681577, 5.640660645, 3.082072075, 3.249229815, 4.492327775,
                  3.488655803, 6.517144589],
    'Observed_Condition': ['Excellent', 'Poor', 'Good', 'Excellent', 'Poor',
                           'Poor', 'Good', 'Good', 'Poor', 'Good', 'Poor', 'Good']
}
df = pd.DataFrame(data)

接下来分三步处理:

  1. 对Observed_Condition为Poor的行抽样:用sample(frac=0.5)抽取该类别的50%数据,frac参数控制抽样比例;如果想抽固定数量,比如3行,换成n=3即可。加上random_state可以让抽样结果固定,方便调试。
  2. 保留其他类别的所有数据:直接筛选出Good和Excellent的行。
  3. 合并两部分数据:用pd.concat()把抽样后的Poor行和其他行合并,ignore_index=True让索引重新排序,避免重复。

代码如下:

# 抽取Poor类别的50%数据
poor_sample = df[df['Observed_Condition'] == 'Poor'].sample(frac=0.5, random_state=42)

# 保留Good和Excellent的所有行
other_rows = df[df['Observed_Condition'].isin(['Good', 'Excellent'])]

# 合并得到最终结果
final_df = pd.concat([poor_sample, other_rows], ignore_index=True)

结果说明

运行后,final_df里会包含:

  • 原数据中Poor类别的3行(原6行的一半)
  • Good和Excellent类别的所有8行

打印final_df会得到类似这样的结果(因为random_state=42,抽样结果固定):

Predicted Observed_Condition
0   3.114821                Poor
1   5.318448                Poor
2   3.249230                Poor
3   4.561075           Excellent
4   5.472004                Good
5   7.048857           Excellent
6   5.640661                Good
7   3.082072                Good
8   4.492328                Good
9   6.517145                Good

如果需要对多个类别分别抽样,也可以用类似的逻辑:分别处理每个需要抽样的组,再和保留组合并即可。

内容的提问来源于stack exchange,提问作者user308827

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:50:08