如何批量生成300个Bernoulli与Poisson随机变量以获取1000万取值
Absolutely you can generate all those 300 variables in one go—no need to manually create each one! This is a common pain point when working with large statistical datasets, and there are straightforward, automated ways to handle it in most popular programming languages. Let’s break down solutions for the two most commonly used tools below:
Python (with NumPy & Pandas)
Python’s numerical libraries make this task trivial. We’ll calculate the sample size per variable (handling the remainder from dividing 10M by 300), then loop to generate all variables at once before exporting to CSV.
import numpy as np import pandas as pd # Define core parameters total_samples = 10_000_000 num_vars = 300 samples_per_var = total_samples // num_vars remainder = total_samples % num_vars # Generate Bernoulli variables (adjust `p` to your desired probability) bernoulli_cols = [] for idx in range(num_vars): # Assign extra sample to first `remainder` columns to hit exactly 10M total current_samples = samples_per_var + 1 if idx < remainder else samples_per_var bern_col = np.random.binomial(n=1, p=0.5, size=current_samples) bernoulli_cols.append(pd.Series(bern_col, name=f"bernoulli_{idx+1}")) # Generate Poisson variables (adjust `lam` to your desired rate parameter) poisson_cols = [] for idx in range(num_vars): current_samples = samples_per_var + 1 if idx < remainder else samples_per_var pois_col = np.random.poisson(lam=2, size=current_samples) poisson_cols.append(pd.Series(pois_col, name=f"poisson_{idx+1}")) # Combine into DataFrames and export to CSV bernoulli_df = pd.concat(bernoulli_cols, axis=1) poisson_df = pd.concat(poisson_cols, axis=1) bernoulli_df.to_csv("bernoulli_distributions.csv", index=False) poisson_df.to_csv("poisson_distributions.csv", index=False)
Quick Python Tips:
- Tweak
p(Bernoulli probability) andlam(Poisson rate) to match your project’s specific requirements. - If you’re tight on memory, generate columns in smaller batches and append them to the CSV incrementally instead of building the full DataFrame upfront.
R (with dplyr & purrr)
R’s functional programming tools eliminate repetitive code for column creation. We’ll use map_dfc to generate and combine columns into a single data frame seamlessly.
library(dplyr) library(purrr) # Define core parameters total_samples <- 10000000 num_vars <- 300 samples_per_var <- total_samples %/% num_vars remainder <- total_samples %% num_vars # Generate Bernoulli variables (adjust `prob` as needed) bernoulli_df <- map_dfc(1:num_vars, function(col_idx) { current_samples <- ifelse(col_idx <= remainder, samples_per_var + 1, samples_per_var) rbinom(n = current_samples, size = 1, prob = 0.5) %>% set_names(paste0("bernoulli_", col_idx)) }) # Generate Poisson variables (adjust `lambda` as needed) poisson_df <- map_dfc(1:num_vars, function(col_idx) { current_samples <- ifelse(col_idx <= remainder, samples_per_var + 1, samples_per_var) rpois(n = current_samples, lambda = 2) %>% set_names(paste0("poisson_", col_idx)) }) # Export to CSV write.csv(bernoulli_df, "bernoulli_distributions.csv", row.names = FALSE) write.csv(poisson_df, "poisson_distributions.csv", row.names = FALSE)
Quick R Tips:
- If you prefer avoiding tidyverse packages, use
base Ralternatives likelapplyand combine results withdo.call(cbind, ...). - For ultra-large datasets, swap
dplyrwithdata.tablefor faster, more memory-efficient operations.
General Best Practices
- Double-check the total sample count: the code above handles the remainder (10M isn’t perfectly divisible by 300) by adding one extra sample to the first 100 columns.
- If your CSV reader struggles with 300 columns, split the output into smaller batches (e.g., 3 files of 100 columns each) by slicing the data frame before exporting.
内容的提问来源于stack exchange,提问作者Random it guy

