如何在Python Pandas DataFrame中生成带概率权重的性别虚拟数据(附R实现代码)
Replicating R's
sample() Gender Data Generation in Pandas No problem at all! Let's replicate exactly what you did in R using Python and Pandas. Here's a straightforward implementation:
First, you'll need to import pandas and numpy (since numpy's random functions handle weighted sampling perfectly):
import pandas as pd import numpy as np
Then, generate the gender sample and wrap it into a DataFrame:
# Generate the weighted sample (matches your R code exactly) gender_sample = np.random.choice(a=["M", "F"], p=[0.6, 0.4], size=100, replace=True) # Convert to a Pandas DataFrame gender_df = pd.DataFrame(gender_sample, columns=["gender"])
Let's map each parameter to your original R code to confirm it's identical:
a=["M", "F"]→ matches R'sx=c("M","F")(the values to sample from)p=[0.6, 0.4]→ matches R'sprob = c(.6, .4)(the weighted probabilities)size=100→ same as R'ssize=100(number of samples to generate)replace=True→ same as R'sreplace=TRUE(allows sampling with replacement, which is required for generating multiple samples from a small value set)
To double-check the distribution aligns with your expected probabilities, run this quick validation:
print(gender_df["gender"].value_counts(normalize=True))
You'll get output close to 60% for "M" and 40% for "F" (tiny random variations are normal each run).
内容的提问来源于stack exchange,提问作者silent_hunter
相关产品推荐
相关产品推荐

