如何使用NumPy或SciPy生成行服从正态分布、列服从均匀分布的二维数组
Hey Stephen! Let's dig into your problem. First, I need to clarify a key point: if you mean that each element in the array follows a normal distribution when viewed row-wise (i.e., every element x_ij ~ Normal for fixed i) and a uniform distribution when viewed column-wise (x_ij ~ Uniform for fixed j), this is theoretically impossible—since a single value can't follow two different distributions at once.
But I suspect you might mean that the empirical distribution of each row's elements approximates a normal distribution, and the empirical distribution of each column's elements approximates a uniform distribution. That makes sense, and we can construct such an array using quantile transformations. Here's how:
Step-by-Step Implementation
1. Generate Initial Normal-Distributed Rows
First, create an array where each row consists of samples from a normal distribution. We'll use the standard normal (mean 0, std 1) for simplicity, but you can adjust parameters as needed:
import numpy as np # Generate 10000 rows, each with 1024 standard normal samples normal_array = np.random.normal(loc=0, scale=1, size=(10000, 1024))
At this point, every row's empirical distribution is approximately normal, but columns are also normally distributed (which doesn't meet your column requirement).
2. Transform Columns to Uniform Distribution
We'll apply a quantile transformation to each column. This converts any distribution to a uniform distribution by mapping each value to its rank-based percentile in the column.
For each column, we:
- Sort the column values
- Assign each value its empirical cumulative distribution function (ECDF) value (i.e., the fraction of values ≤ it in the column)
- Replace the original values with these ECDF values
Here's the code to do this efficiently:
def quantile_transform_column(col): # Get sorted indices and sorted values sorted_indices = np.argsort(col) sorted_vals = col[sorted_indices] # Calculate ECDF values (percentiles) ecdf = np.linspace(1/len(col), 1, len(col)) # Handle duplicates by assigning the average ECDF for tied values unique_vals, idx = np.unique(sorted_vals, return_inverse=True) avg_ecdf = np.bincount(idx, weights=ecdf) / np.bincount(idx) # Map back to the original column order transformed_col = np.empty_like(col) transformed_col[sorted_indices] = avg_ecdf[idx] return transformed_col # Apply the transformation to every column result_array = np.apply_along_axis(quantile_transform_column, axis=0, arr=normal_array)
3. Verify the Results
You can check the empirical distributions to confirm:
- For any row, plot a histogram of its elements—it should closely resemble a normal curve (note: since we used column-specific transformations, the row's distribution won't be perfectly normal, but the deviation will be minimal with 1024 elements per row)
- For any column, plot a histogram—it should show a flat, uniform distribution across the [0,1] range
Optional: Adjust Uniform Distribution Range
If you need columns to follow a uniform distribution over [a, b] instead of [0,1], simply scale the result:
a = 2 b = 5 result_array_scaled = a + (b - a) * result_array
Key Notes
- If you strictly need rows to follow a theoretical normal distribution while columns follow a theoretical uniform distribution, this is mathematically impossible (as each element's marginal distribution can't be both normal and uniform). The approach above gives the closest practical approximation for empirical distributions.
- The row-wise empirical distribution remains close to normal because quantile transformations are monotonic—they preserve the relative order of values, so the overall shape of each row's distribution stays similar to the original normal.
内容的提问来源于stack exchange,提问作者Stephen Wong

