You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

分布式进程崩溃后,如何基于MySQL与ZooKeeper恢复任务信息?

Process Crash Recovery in Distributed Systems (MySQL + ZooKeeper Setup)

Great question—this is a super common pain point in distributed systems, and your stack (MySQL + ZooKeeper) gives you all the tools you need to build a solid recovery flow. Let’s break this down into actionable steps that play to each tool’s strengths.

Core Approach

The key is to:

  1. Track which processes are alive (ZooKeeper’s bread and butter)
  2. Maintain clear, transactional state for tasks in MySQL
  3. Safely claim and resume crashed processes’ tasks without duplicates or data loss

Step-by-Step Implementation

1. Use ZooKeeper for Process Liveness Tracking

Every time a process starts up, have it create a ephemeral node in ZooKeeper (e.g., /workers/worker-<unique-id>). Ephemeral nodes are automatically deleted when the process disconnects (crash, network split, etc.).

  • Store a small payload in the node: the worker’s unique ID (could be a UUID or sequential ID from ZK) and maybe its current load.
  • Set up a Watcher on the /workers path (either in a dedicated monitoring service or other worker processes) to detect when a node disappears. This is your first trigger for recovery.

2. Enhance MySQL Task Table for State Tracking

Update your task table to include these critical fields (on top of your existing task data):

  • worker_id: The unique ID of the process currently handling the task
  • task_status: Enum with values like PENDING, RUNNING, RECOVERING, COMPLETED, FAILED
  • heartbeat_ts: Timestamp updated by the worker every N seconds (e.g., 10s) to prove it’s alive
  • progress_details: Optional JSON field to store checkpoint data (e.g., "last_processed_record_id": 1234) for resumable tasks

Example schema snippet:

ALTER TABLE tasks 
ADD COLUMN worker_id VARCHAR(64),
ADD COLUMN task_status ENUM('PENDING', 'RUNNING', 'RECOVERING', 'COMPLETED', 'FAILED') DEFAULT 'PENDING',
ADD COLUMN heartbeat_ts DATETIME,
ADD COLUMN progress_details JSON;

3. Trigger Recovery When a Process Crashes

You have two complementary ways to detect a crashed worker:

  • ZooKeeper Watcher Trigger: When a worker’s ephemeral node is deleted, the watcher immediately marks all tasks assigned to that worker as RECOVERING in MySQL (use a transaction to ensure atomicity).
  • MySQL Heartbeat Check: Run a periodic job (e.g., every 30s) that scans for RUNNING tasks where heartbeat_ts is older than your threshold (e.g., 20s). Mark these as RECOVERING—this acts as a fallback if ZooKeeper has a temporary blip.

4. New Worker Claims and Resumes Tasks

When a new process starts up:

  1. Register itself in ZooKeeper (create an ephemeral node, get its unique worker_id).
  2. Query MySQL for tasks marked RECOVERING that were assigned to the crashed worker. Use an optimistic lock to claim tasks safely:
    UPDATE tasks 
    SET worker_id = '<new-worker-id>', task_status = 'RUNNING', heartbeat_ts = NOW()
    WHERE worker_id = '<crashed-worker-id>' AND task_status = 'RECOVERING';
    
    This ensures only one worker can claim each task (MySQL’s row-level locking prevents race conditions).
  3. For each claimed task:
    • If it’s a resumable task, load the progress_details and pick up from the last checkpoint.
    • If it’s a non-resumable task (e.g., a one-time API call), verify it wasn’t already completed (check logs or an is_executed flag) before re-running.
  4. Start sending regular heartbeats to update heartbeat_ts for all active tasks.

5. Guard Against Edge Cases

  • Avoid Duplicate Execution: Always use the optimistic lock update when claiming tasks—never just select and update separately (that’s a race condition waiting to happen).
  • Handle Partial Failures: If a worker crashes mid-task, make sure your task logic is idempotent (can be run multiple times without side effects) or uses checkpointing to resume from where it left off.
  • ZooKeeper Split Brain: Since ZooKeeper uses quorum consensus, it’s resistant to split brain, but your recovery logic should rely on the ephemeral node deletion as the source of truth for worker liveness.

Bonus Optimizations

  • Task Sharding: Assign tasks to workers using a consistent hash on task_id and worker_id—this makes recovery faster because new workers only need to scan a subset of tasks.
  • Recovery Logging: Add a task_recovery_log table to track every recovery attempt (worker ID, timestamp, outcome) for debugging.
  • Auto-Scaling Integration: Tie the recovery trigger to your auto-scaler—if a worker crashes, spin up a new one automatically and kick off the task claim process.

内容的提问来源于stack exchange,提问作者hehe

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 21:13:13