You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Pandas DataFrame中随机将指定列的K个值设置为空?

第一步:先修正random_rows函数的错误

你现有函数里误拿列数(df.shape[1])当做行数计算,会导致索引范围匹配错误,修正后的函数如下:

import random
import pandas as pd

def random_rows(K=2, df):
    # 获取DataFrame的行数
    row_length = df.shape[0]
    row_indexes = list(range(row_length))
    if row_length < K:
        K = row_length
    selected_row_indexes = random.sample(row_indexes, K)
    return selected_row_indexes

第二步:修改指定行的A列值

拿到随机选中的行索引后,直接用pandas的loc索引器定位到对应行和A列,赋值为空即可,完整运行示例如下:

# 构造测试DataFrame
df = pd.DataFrame( { 'A': [1,2,3,4],
                   'B': [10,20,30,40],
                   'C': [20,40,60,80]
                  })
# 随机选K=2行
selected_indexes = random_rows(K=2, df)
# 将选中行的A列设为空字符串,如果需要pandas标准缺失值可替换为pd.NA
df.loc[selected_indexes, 'A'] = ''
# 输出验证结果
print(df)

补充说明

  • 如果你的DataFrame索引不是默认的0起始整数索引,可改用iloc按位置索引修改,代码调整为df.iloc[selected_indexes, df.columns.get_loc('A')] = ''即可
  • 赋值为pd.NA更符合pandas缺失值规范,后续做缺失值统计、过滤等操作兼容性更好

内容的提问来源于stack exchange,提问作者marlon

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.01 13:15:03