You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas替换Age区间为对应随机数时同区间值相同的代码修正

问题原因

你代码的核心问题出在替换字典的构造逻辑:randint() 属于即时执行函数,在你构造{'65+': randint(65,100), ...}这个字典时,四个randint()就已经分别运行一次,生成了四个固定的随机值作为替换目标,后续replace操作只会把所有匹配到的同分组值统一替换为对应固定值,不会每行重新生成随机数。

修正方案

方案1:向量化批量生成(推荐,适配大型数据集,性能更高)

该方案避免了逐行遍历的开销,适合你提到的大型数据集场景:

import numpy as np

# 定义各年龄区间的取值范围
age_range_map = {
    '65+': (65, 100),
    '16-25': (16, 25),
    '26-39': (26, 39),
    '40-64': (40, 64)
}

# 按分组批量生成对应长度的随机数赋值
for age_bin, (low, high) in age_range_map.items():
    match_mask = df['Age'] == age_bin
    df.loc[match_mask, 'Age'] = np.random.randint(low, high + 1, size=match_mask.sum())

方案2:apply逐行处理(写法简洁,适合中小数据集)

该方案逻辑易读,代码量更小:

import random

def bin_to_random_age(age_bin):
    low, high = {
        '65+': (65, 100),
        '16-25': (16, 25),
        '26-39': (26, 39),
        '40-64': (40, 64)
    }[age_bin]
    return random.randint(low, high)

df['Age'] = df['Age'].apply(bin_to_random_age)

注意:请根据你实际的列名调整代码中的Age字段,保持和数据集列名一致即可。


内容的提问来源于stack exchange,提问作者Ahmed JAÏEM

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.04 07:54:04