You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从核密度估计(KDE)生成匹配原始百分比收益率的随机样本?

Let's work through this issue to get your KDE-generated samples matching your percentage-based return data properly.

1. Align Your Return Scale First

Right now, your returns = data.pct_change(1) produces decimal returns (e.g., a 2% monthly gain shows up as 0.02). If you want your samples to be in percentage form (like 2 instead of 0.02), you need to scale these returns explicitly:

# Convert decimal returns to percentage values (so 0.02 becomes 2)
returns = data.pct_change(1) * 100
returns = returns.iloc[1:]  # Drop the first row with NaN

This ensures your KDE models are fitted to the exact scale you expect for your simulated samples.

2. Add Proper Sample Generation from KDE Models

Your code fits the KDE models correctly, but you're missing the step to generate random samples from those models. Here's how to pull 100 samples per stock:

import numpy as np
from sklearn.model_selection import GridSearchCV
from sklearn.neighbors import KernelDensity

# Your existing KDE fitting code (now using scaled percentage returns)
params = {'bandwidth': np.logspace(-1, 1, 20)}
grid = GridSearchCV(KernelDensity(), params)
models = returns.apply(lambda x: grid.fit(x.values.reshape(-1, 1)).best_estimator_)

# Generate 100 simulated percentage returns per stock
num_sim_samples = 100
simulated_returns = models.apply(lambda kde: kde.sample(num_sim_samples).flatten())

3. Diagnose the 20-200 Sample Range Problem

If you're still seeing samples in the 20-200 range when your actual returns are reasonable (e.g., -10% to 10% for DJIA stocks), check these key points:

  • Validate your raw returns: Run print(returns.describe()) to see the min/max/mean of your real returns. If the original data already has returns in the 20-200 range, your data might not be price data (or it's extremely volatile, which is unusual for DJIA components).
  • Tweak KDE bandwidth: The np.logspace(-1, 1, 20) range gives bandwidths from 0.1 to 10. For percentage returns (typically -10 to 10), try a smaller range like np.logspace(-2, 0, 20) (0.01 to 1) to avoid over-smoothing the distribution, which can cause extreme sample values.
  • Double-check data shape: The x.values.reshape(-1, 1) is correct for scikit-learn's KernelDensity, but confirm each column in returns is a 1D array of monthly returns for a single stock.

4. Quick Monte Carlo Path Example

Once you have valid simulated percentage returns, you can generate 100 paths per stock (e.g., 60 steps each) like this:

# Grab the last observed price for each stock
last_prices = data.iloc[-1]
num_path_steps = 60
monte_carlo_paths = {}

for stock, kde_model in models.items():
    # Generate returns for each step in each path
    step_returns = kde_model.sample((num_path_steps, num_sim_samples)).reshape(num_path_steps, num_sim_samples)
    # Convert back to decimal returns for price calculations
    decimal_step_returns = step_returns / 100
    # Build price paths from cumulative returns
    price_paths = last_prices[stock] * (1 + decimal_step_returns).cumprod(axis=0)
    monte_carlo_paths[stock] = price_paths

内容的提问来源于stack exchange,提问作者Karl Hunt

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 04:08:41