You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

关于Nature 2015经典DQN论文的训练方式问询:单DQN还是49个DQN?

Clarification on DQN Training Setup in the 2015 Nature Paper

Great question—this is a common point of confusion when revisiting the foundational DQN work! Let's break this down clearly:

Core Answer

The paper Human level control through deep reinforcement learning (2015, Nature) uses 49 independently trained neural networks, one initialized from scratch for each of the 49 Atari games. A single model was NOT used across all games.

Context & Evidence

While the paper stresses that the same algorithm, network architecture, and hyperparameters were applied uniformly to every game, the experimental methodology (and supplementary materials) make it explicit that each game got its own dedicated training run:

  • Each network started with random initial weights, with no parameter sharing, transfer learning, or multi-task training between games.
  • This choice made sense for 2015: multi-task reinforcement learning that could handle the diverse mechanics of Atari games was still in its early stages, and independent training let each model fully adapt to the unique state space, reward signals, and gameplay of individual games.

Why This Difference Matters

  • Independent per-game training: Each model optimizes exclusively for one game's dynamics, which directly contributed to the paper's landmark results of matching or exceeding human-level performance across most titles.
  • Single multi-game model: This would require a single network to learn wildly different behaviors (e.g., navigating a maze in Pac-Man vs. controlling a paddle in Pong) with shared weights—a challenge that became more feasible in later RL work (like Rainbow or multi-task frameworks) but wasn't the approach here.

As a side note, checking the original DeepMind code releases from that era confirms this: the training pipeline was structured to run per-game, with separate checkpoints, logs, and model artifacts for each of the 49 games.

内容的提问来源于stack exchange,提问作者T.L

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:24:44