You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python Pandas DataFrame中生成带概率权重的性别虚拟数据(附R实现代码)

Replicating R's sample() Gender Data Generation in Pandas

No problem at all! Let's replicate exactly what you did in R using Python and Pandas. Here's a straightforward implementation:

First, you'll need to import pandas and numpy (since numpy's random functions handle weighted sampling perfectly):

import pandas as pd
import numpy as np

Then, generate the gender sample and wrap it into a DataFrame:

# Generate the weighted sample (matches your R code exactly)
gender_sample = np.random.choice(a=["M", "F"], p=[0.6, 0.4], size=100, replace=True)

# Convert to a Pandas DataFrame
gender_df = pd.DataFrame(gender_sample, columns=["gender"])

Let's map each parameter to your original R code to confirm it's identical:

  • a=["M", "F"] → matches R's x=c("M","F") (the values to sample from)
  • p=[0.6, 0.4] → matches R's prob = c(.6, .4) (the weighted probabilities)
  • size=100 → same as R's size=100 (number of samples to generate)
  • replace=True → same as R's replace=TRUE (allows sampling with replacement, which is required for generating multiple samples from a small value set)

To double-check the distribution aligns with your expected probabilities, run this quick validation:

print(gender_df["gender"].value_counts(normalize=True))

You'll get output close to 60% for "M" and 40% for "F" (tiny random variations are normal each run).

内容的提问来源于stack exchange,提问作者silent_hunter

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.27 15:07:32