如何从核密度估计(KDE)生成匹配原始百分比收益率的随机样本?
Let's work through this issue to get your KDE-generated samples matching your percentage-based return data properly.
1. Align Your Return Scale First
Right now, your returns = data.pct_change(1) produces decimal returns (e.g., a 2% monthly gain shows up as 0.02). If you want your samples to be in percentage form (like 2 instead of 0.02), you need to scale these returns explicitly:
# Convert decimal returns to percentage values (so 0.02 becomes 2) returns = data.pct_change(1) * 100 returns = returns.iloc[1:] # Drop the first row with NaN
This ensures your KDE models are fitted to the exact scale you expect for your simulated samples.
2. Add Proper Sample Generation from KDE Models
Your code fits the KDE models correctly, but you're missing the step to generate random samples from those models. Here's how to pull 100 samples per stock:
import numpy as np from sklearn.model_selection import GridSearchCV from sklearn.neighbors import KernelDensity # Your existing KDE fitting code (now using scaled percentage returns) params = {'bandwidth': np.logspace(-1, 1, 20)} grid = GridSearchCV(KernelDensity(), params) models = returns.apply(lambda x: grid.fit(x.values.reshape(-1, 1)).best_estimator_) # Generate 100 simulated percentage returns per stock num_sim_samples = 100 simulated_returns = models.apply(lambda kde: kde.sample(num_sim_samples).flatten())
3. Diagnose the 20-200 Sample Range Problem
If you're still seeing samples in the 20-200 range when your actual returns are reasonable (e.g., -10% to 10% for DJIA stocks), check these key points:
- Validate your raw returns: Run
print(returns.describe())to see the min/max/mean of your real returns. If the original data already has returns in the 20-200 range, yourdatamight not be price data (or it's extremely volatile, which is unusual for DJIA components). - Tweak KDE bandwidth: The
np.logspace(-1, 1, 20)range gives bandwidths from 0.1 to 10. For percentage returns (typically -10 to 10), try a smaller range likenp.logspace(-2, 0, 20)(0.01 to 1) to avoid over-smoothing the distribution, which can cause extreme sample values. - Double-check data shape: The
x.values.reshape(-1, 1)is correct for scikit-learn'sKernelDensity, but confirm each column inreturnsis a 1D array of monthly returns for a single stock.
4. Quick Monte Carlo Path Example
Once you have valid simulated percentage returns, you can generate 100 paths per stock (e.g., 60 steps each) like this:
# Grab the last observed price for each stock last_prices = data.iloc[-1] num_path_steps = 60 monte_carlo_paths = {} for stock, kde_model in models.items(): # Generate returns for each step in each path step_returns = kde_model.sample((num_path_steps, num_sim_samples)).reshape(num_path_steps, num_sim_samples) # Convert back to decimal returns for price calculations decimal_step_returns = step_returns / 100 # Build price paths from cumulative returns price_paths = last_prices[stock] * (1 + decimal_step_returns).cumprod(axis=0) monte_carlo_paths[stock] = price_paths
内容的提问来源于stack exchange,提问作者Karl Hunt

