You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中基于多索引DataFrame高效生成随机数据的方法

嘿,这个问题我之前处理类似数据集时也踩过循环的坑——手动遍历多索引不仅代码啰嗦,数据量大的时候简直慢到让人崩溃。咱们直接用Pandas+NumPy的矢量化操作来搞定,既高效又简洁!

先把你给出的统计数据整理成清晰的表格方便参考:

原始统计数据(多索引DataFrame)

LeaderExpense_Typemeanstdcount
Leader1Airfare1979.6842192731.6297671358
Leader1Booking Fees118.994538270.0073901179
Leader1Conference/Seminars1553.8309231319.29594665
Leader1Hotel1656.6436582104.7210931405
Leader1Meals435.665122676.7058571476
Leader1Mileage213.785046284.908031979
Leader1Taxi/Uber308.530724380.2889641422
Leader2Airfare1730.1969112334.688155628
Leader2Booking Fees112.020556573.407269576
Leader2Conference/Seminars1647.5765001154.32058480
Leader2Hotel1693.0803561953.552474618
Leader2Meals574.228548844.997595620
Leader2Mileage215.898798291.231331466
Leader2Taxi/Uber298.655852340.926518569

高效实现方案

方案1:简洁版(Pandas Apply + Explode)

这个方案代码最直观,适合大多数场景,比手动循环高效N倍:

import pandas as pd
import numpy as np

# 假设你的多索引统计DataFrame名为 df_stats
# 先重置索引,方便后续处理(最后可恢复)
df_flat = df_stats.reset_index()

# 定义生成对应数量随机数的函数
def generate_expenses(row):
    # 生成count个符合正态分布的随机数
    return np.random.normal(loc=row['mean'], scale=row['std'], size=row['count']).tolist()

# 给每行生成随机数列表
df_flat['Expense_Amount'] = df_flat.apply(generate_expenses, axis=1)

# 把列表展开成单独的行,同时丢弃原统计列
df_random = df_flat.explode('Expense_Amount').drop(['mean', 'std', 'count'], axis=1)

# 可选:恢复多索引结构
df_random = df_random.set_index(['Leader', 'Expense_Type'])

方案2:极致高效版(全矢量化NumPy操作)

如果你的实际数据量特别大(比如成百上千个Leader/Expense_Type组合),上面的apply还有一点点Python级别的开销,推荐用全NumPy矢量化操作,性能拉满:

import pandas as pd
import numpy as np

df_flat = df_stats.reset_index()

# 提取统计数据为NumPy数组,方便批量处理
means = df_flat['mean'].values
stds = df_flat['std'].values
counts = df_flat['count'].values

# 生成重复的分组索引,用来匹配每个随机数对应的Leader/Expense_Type
group_indices = np.repeat(np.arange(len(df_flat)), counts)

# 一次性生成所有随机数,用concatenate合并各组结果
all_expenses = np.concatenate([
    np.random.normal(loc=m, scale=s, size=c) 
    for m, s, c in zip(means, stds, counts)
])

# 构造最终的DataFrame
df_random = pd.DataFrame({
    'Leader': df_flat['Leader'].values[group_indices],
    'Expense_Type': df_flat['Expense_Type'].values[group_indices],
    'Expense_Amount': all_expenses
}).set_index(['Leader', 'Expense_Type'])

额外提示

  • 如果你需要非正态分布的随机数,只需要把np.random.normal换成对应的函数(比如均匀分布用np.random.uniform,泊松分布用np.random.poisson等)
  • 可以添加np.random.seed(123)来固定随机种子,保证结果可复现
  • 最终的df_random就是你需要的:每个Leader+Expense_Type组合下,生成了count条符合对应均值和标准差的随机费用数据

内容的提问来源于stack exchange,提问作者Agarp

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 03:59:49