关于Nature 2015经典DQN论文的训练方式问询:单DQN还是49个DQN?
Great question—this is a common point of confusion when revisiting the foundational DQN work! Let's break this down clearly:
Core Answer
The paper Human level control through deep reinforcement learning (2015, Nature) uses 49 independently trained neural networks, one initialized from scratch for each of the 49 Atari games. A single model was NOT used across all games.
Context & Evidence
While the paper stresses that the same algorithm, network architecture, and hyperparameters were applied uniformly to every game, the experimental methodology (and supplementary materials) make it explicit that each game got its own dedicated training run:
- Each network started with random initial weights, with no parameter sharing, transfer learning, or multi-task training between games.
- This choice made sense for 2015: multi-task reinforcement learning that could handle the diverse mechanics of Atari games was still in its early stages, and independent training let each model fully adapt to the unique state space, reward signals, and gameplay of individual games.
Why This Difference Matters
- Independent per-game training: Each model optimizes exclusively for one game's dynamics, which directly contributed to the paper's landmark results of matching or exceeding human-level performance across most titles.
- Single multi-game model: This would require a single network to learn wildly different behaviors (e.g., navigating a maze in Pac-Man vs. controlling a paddle in Pong) with shared weights—a challenge that became more feasible in later RL work (like Rainbow or multi-task frameworks) but wasn't the approach here.
As a side note, checking the original DeepMind code releases from that era confirms this: the training pipeline was structured to run per-game, with separate checkpoints, logs, and model artifacts for each of the 49 games.
内容的提问来源于stack exchange,提问作者T.L

