You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何为DataFrame列随机替换20%内容为2-10个随机数量的NaN?

实现DataFrame按指定规则随机添加NaN

需求说明

需要对Pandas DataFrame的列进行随机NaN替换,满足两个核心条件:

  1. 最终每列的NaN占比约为20%
  2. 每次操作添加的NaN数量在2-10之间随机,分批次完成替换(类似示例中首次加3个、第二次加1个的效果)

代码实现

import pandas as pd
import numpy as np

def add_random_nans(df, target_pct=0.2, min_batch=2, max_batch=10):
    for col in df.columns:
        total_rows = len(df)
        # 计算该列需要达到的目标NaN总数
        target_total = int(total_rows * target_pct)
        # 当前已有的NaN数量
        current_nans = df[col].isna().sum()
        # 还需要添加的NaN数量
        need_add = target_total - current_nans
        
        if need_add <= 0:
            continue
        
        # 分批次添加NaN
        while need_add > 0:
            # 本次批次的NaN数量,取[min_batch, max_batch]和剩余量的最小值
            batch_size = np.random.randint(min_batch, min(max_batch, need_add) + 1)
            # 获取当前列非NaN的行索引
            valid_indices = df[col].dropna().index
            # 随机选择batch_size个索引进行替换
            selected = np.random.choice(valid_indices, size=batch_size, replace=False)
            df.loc[selected, col] = np.nan
            # 更新剩余需要添加的数量
            need_add -= batch_size
    return df

# 示例使用
if __name__ == "__main__":
    # 创建测试用DataFrame
    test_df = pd.DataFrame({
        'Col1': range(1, 101),
        'Col2': range(101, 201),
        'Col3': range(201, 301)
    })
    
    # 执行随机添加NaN操作
    result_df = add_random_nans(test_df)
    
    # 输出结果及NaN占比
    print("处理后DataFrame:")
    print(result_df.head(15))
    print("\n各列NaN占比:")
    print(result_df.isna().mean().round(2))

代码逻辑解释

  1. 目标计算:针对每列,先根据总行数的20%计算需要达到的NaN总数,减去已有NaN数量得到待添加量
  2. 批次控制:循环生成2-10之间的随机批次大小(若剩余待添加量不足10,则取剩余量),避免一次性添加过多NaN
  3. 随机选择:每次从列的非NaN行中随机挑选对应数量的索引,替换为NaN,直到达到目标占比

这种方式既保证了整体NaN占比符合要求,又实现了分批次随机添加的效果,完全匹配需求中的示例表现。

内容的提问来源于stack exchange,提问作者Ahmad Aburoman

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.01 22:10:35