自组织映射(SOM/Kohonen Map)回归器效果异常极差的问题排查:数据缩放与模型适配性问询
Let's walk through the possible issues causing your extremely low R² scores (especially the negative test set scores) and break down whether this is due to data handling mistakes or inherent model limitations.
First: Checking Your Data Scaling Workflow
Your overall scaling approach is correct in theory:
- Splitting data before scaling to avoid leakage ✔️
- Using separate scalers for features and target ✔️
- Applying only the training-set fit to test data ✔️
- Inversing predictions back to original scale before calculating R² ✔️
That said, there's one minor code quirk to watch for: you're overwriting the scaled y_train variable with its inverse-transformed version later in the loop. While this doesn't break the R² calculation (since you're comparing inverse-transformed predictions to inverse-transformed training labels), it's better to rename variables for clarity (e.g., y_train_scaled for the scaled version, y_train_original for the inverse-transformed one) to avoid confusion down the line.
Likely Culprits: SOM Parameter & Configuration Issues
The extreme negative test R² suggests your model isn't learning any meaningful patterns from the data—instead, it's producing predictions that are worse than just guessing the mean of the target variable. Here are the key parameters to adjust:
1. Map Size Calculation
You're using the test set size to calculate the map size:
map_size= int(5* math.sqrt(X_test.shape[0])) #vesanto
Vesanto's formula is intended for the training set size, not the test set. If your test set is much smaller than the training set, this will create an overly small map that can't capture the training data's distribution. Fix this with:
map_size = int(5 * math.sqrt(X_train.shape[0]))
You can also experiment with smaller multipliers (e.g., 2 instead of 5) to avoid an overly large map that leads to sparse neurons and poor generalization.
2. Learning Rate & Iteration Settings
Your current learning rate (start=0.5, end=0.05) is quite high for SOMs—this can cause weight oscillations and prevent the model from converging to a stable topological map. Try:
- Lowering the starting learning rate to 0.1 and ending rate to 0.01
- Separating unsupervised and supervised iteration counts: SOMs need sufficient unsupervised training to map the feature space first, then supervised fine-tuning for regression. For example:
som = susi.SOMRegressor( n_iter_unsupervised=2000, # More unsupervised iterations to learn topology n_iter_supervised=500, # Less supervised fine-tuning learning_rate_start=0.1, learning_rate_end=0.01, # ... other parameters )
3. Neighborhood Mode
Linear neighborhood mode might not be ideal for all datasets. Try switching to Gaussian neighborhood mode for unsupervised training, which decays the neighborhood influence more smoothly as iterations progress:
neighborhood_mode_unsupervised="gaussian"
Also, check if you can set a decaying neighborhood radius (some SOM implementations let you define neighborhood_radius_start and neighborhood_radius_end—verify the susi docs for this).
Data & Model Suitability Checks
1. Is Your Data Predictable?
Before blaming the SOM, test with a simple baseline model like linear regression or a decision tree. If these models also produce terrible R² scores, your features likely have little to no predictive power for the target variable.
- Calculate feature-target correlations with
X.corrwith(y)to see if any features are strongly correlated with your target. - Check for distribution shifts between training and test sets (e.g., plot histograms of
y_trainandy_test—if they look drastically different, the model can't generalize).
2. Is SOM the Right Tool for This Task?
SOMs excel at high-dimensional data with topological structure, but they're not the best choice for all regression problems—especially if your data is time-series (your currency variable suggests this). Time-series data has temporal dependencies that SOMs don't explicitly model. If this is time-series data, consider trying specialized models like LSTMs, ARIMA, or even XGBoost with lag features first.
Quick Debugging Steps
- Inspect Predictions: Print the distribution of
y_predand compare it toy_test. If predictions have an absurdly large range or are completely disconnected from the test data, your scaling or model parameters are off. - Check Convergence: Track training R² across iterations—if it doesn't improve over time, your model isn't learning, and you need to adjust learning rates or iteration counts.
- Validate Scaling: Print the min/max of
y_train_scaled(should be 0-1) andy_predbefore inverse transformation (also should be 0-1) to confirm scaling is working as intended.
内容的提问来源于stack exchange,提问作者LGR

