关于采用非梯度优化方法训练易梯度爆炸系统的文献问询
Absolutely, there’s a robust body of research focused on using non-gradient-based optimization methods to train systems like recurrent neural networks (RNNs) that are prone to gradient explosion. Your intuition about gradient explosion creating sharp ridges in the loss landscape—causing gradient-based methods to overshoot minima by updating parameters proportionally to extreme gradients—is perfectly aligned with why these non-gradient approaches are valuable.
Here are key areas of research and notable work in this space:
Evolutionary Algorithms (EAs)
Genetic algorithms (GAs) are a classic choice for gradient-free RNN training. Early work like "Evolutionary Optimization of Recurrent Neural Networks for Time Series Prediction" showed that GAs can navigate loss landscapes without gradient signals, avoiding divergence from exploding gradients. By evolving populations of RNN weights via selection, crossover, and mutation, GAs don’t follow the steep ridges that gradient descent often gets stuck on or overshoots. Later extensions, like neuroevolution techniques, have refined this for more complex sequence tasks.Particle Swarm Optimization (PSO)
PSO’s swarm-based search strategy is well-suited for RNNs with gradient stability issues. Studies such as "Training Recurrent Neural Networks with Particle Swarm Optimization" demonstrated that PSO can converge to viable minima even when gradient-based methods fail. Instead of following gradient directions, particles (each representing a set of RNN weights) adjust their positions based on their own best performance and the swarm’s collective best, which makes them far less likely to be thrown off by sharp ridges.Simulated Annealing (SA)
Inspired by metallurgical cooling processes, simulated annealing has been applied to RNN training to mitigate gradient explosion problems. Work like "Simulated Annealing for Training Recurrent Neural Networks" leverages SA’s ability to temporarily accept worse solutions early in the search to escape local minima and avoid being trapped by the steep, gradient-explosion-induced ridges. As the algorithm’s "temperature" decreases, it converges to more stable, optimal parameter sets.Modern Gradient-Free Variants
More recent research has explored hybrid approaches (combining non-gradient methods with gradient-based tweaks) and specialized gradient-free optimizers like CMA-ES (Covariance Matrix Adaptation Evolution Strategy), which have shown promise for training deep RNNs and transformers in scenarios where gradient explosion is a persistent issue.
Your core point about gradient-based methods overshooting minima due to proportional updates to extreme gradients is exactly why non-gradient methods shine here. By abandoning gradient signals entirely, these approaches use heuristic exploration to navigate the loss landscape, avoiding the sudden, massive parameter jumps that come with exploding gradients.
内容的提问来源于stack exchange,提问作者LiKao

