You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

请求指引TensorFlow分布式训练源码与Worker-PS架构位置

Hey there! I totally get how overwhelming the TensorFlow codebase can be when you're trying to unpack distributed training—let me walk you through the key directories and source files to make this easier.

First, High-Level Repository Structure

TensorFlow splits its distributed logic across core C++ runtime (low-level communication/coordination) and Python APIs (user-facing strategies and utilities). Here’s the lay of the land:

  • tensorflow/core/: The heart of TensorFlow’s C++ implementation, where most distributed runtime logic lives.
  • tensorflow/python/distribute/: Home to modern distributed training APIs (like ParameterServerStrategy, MirroredStrategy). This is where you’ll find the Python wrappers that abstract Worker/PS interactions.
  • tensorflow/python/training/: Contains legacy distributed training utilities (still relevant if you’re looking at older PS-based workflows).

Key Directories for Worker & PS Logic

1. Core Distributed Runtime (C++ Layer)

This is where the actual Worker/PS communication, service definitions, and parameter management happen:

  • tensorflow/core/distributed_runtime/: The main hub for distributed execution:
    • server_lib.h/server_lib.cc: Defines the underlying Server class that powers both Worker and PS nodes. Every node in your cluster is an instance of this, initialized with a role (worker/ps/master).
    • worker_service.h/worker_service.cc: Implements the RPC service that Worker nodes expose. This handles tasks like executing graph ops, requesting parameters from PS nodes, and sending gradients back.
    • parameter_server/: A dedicated subdirectory for PS node logic:
      • parameter_server.h/parameter_server.cc: Implements parameter storage, gradient application, and concurrency control for PS nodes.
      • parameter_server_service.h/parameter_server_service.cc: The RPC service interface for PS nodes, handling read/write requests from Workers.
    • master_service.h/master_service.cc: The master coordinator that schedules tasks across Workers and PS nodes, managing the overall distributed graph execution.

2. Python Layer Wrappers & APIs

If you’re starting from user-facing code, these files will lead you to the core logic:

  • tensorflow/python/distribute/parameter_server_strategy.py: The modern API for PS-based distributed training. This wraps the C++ runtime logic, handling Worker/PS initialization, variable placement, and gradient aggregation.
  • tensorflow/python/training/server_lib.py: Defines the Python tf.train.Server class—this is what you use to start Worker/PS nodes in legacy workflows. It calls into the C++ Server implementation under the hood.
  • tensorflow/python/training/device_setter.py: Contains replica_device_setter, a utility that maps model variables to PS nodes (critical for legacy PS setups to ensure variables live on PS instead of Workers).

Worker & PS Architecture Overview

To tie this all together:

  1. Cluster Initialization: You start Worker and PS nodes using tf.train.Server (or ParameterServerStrategy’s cluster setup). Each node registers its RPC service (WorkerService for Workers, ParameterServerService for PS).
  2. Task Scheduling: The master service (or strategy’s internal coordinator) splits the computation graph: variables are placed on PS nodes, while compute-heavy ops (forward/backward passes) are assigned to Workers.
  3. Execution Flow:
    • Workers send RPC requests to PS nodes to fetch current parameter values.
    • Workers compute gradients using local batch data.
    • Workers send gradients back to PS nodes via RPC.
    • PS nodes apply gradients to update parameters, handling concurrent requests with locking/atomic operations.

Pro Tip for Navigation

Start with the Python API you’re familiar with (e.g., ParameterServerStrategy), then trace the calls down into the C++ runtime. For example, look for _create_distributed_runtime in parameter_server_strategy.py—this will lead you straight to the core distributed runtime code in tensorflow/core/distributed_runtime/.

内容的提问来源于stack exchange,提问作者Huzo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 07:13:46