Nomad调度器节点选择异常咨询:多单实例Job仅调度至首节点
Alright, let's break down why all your single-instance jobs are landing on the first node and how to fix this. The behavior you're seeing points to either default scheduler behavior combined with node priority settings, or subtle cluster node issues that make the first node the only "preferred" target.
First, Verify Your Cluster Node Health & Configuration
Before tweaking job settings, let's rule out basic cluster problems:
- Check node status: Run
nomad node statusto confirm nodes 2 and 3 are marked asreadyandschedulable. If either is indrainingordownstate, that’s why jobs aren’t landing there. - Inspect node resources: Use
nomad node status <node-id>for nodes 2 and 3 to check available CPU/memory. Even "simple" Java jobs need baseline resources—if these nodes are maxed out, Nomad will skip them. - Validate Java driver setup: Ensure the Java driver is properly installed and enabled on all three nodes. Run
nomad node status <node-id> -verboseand check theDriverssection—ifjavaisn’t listed for nodes 2/3, that’s a critical blocker.
Fix the Scheduler Behavior
Since you confirmed nodes 2/3 work when using distinct_host with count=2, the issue is that Nomad’s default scheduler isn’t spreading single-instance jobs across nodes. Here’s how to force even distribution:
1. Add a spread Constraint to Your Jobs
The spread directive tells Nomad to prioritize placing tasks on nodes with the fewest allocations. Add this to your job’s group block to override default packing behavior:
job "test" { datacenters = ["dc1"] type = "service" group "test" { count = 1 # Force scheduler to spread instances across all nodes spread { attribute = "${node.unique.name}" weight = 100 } task "test" { driver = "java" config { # Your existing Java config here jar_path = "/path/to/your/service.jar" # ... other driver configurations } } } }
The weight=100 ensures this spread rule takes high priority over Nomad’s default behavior.
2. Check Node Weights
Nomad uses node weights (default 100) to prioritize scheduling targets. If your first node has a higher weight (e.g., 200), Nomad will favor it for all new jobs. Verify weights with:
nomad node status -verbose
If node 1’s weight is elevated, reset it to match the others using:
nomad node update -weight 100 <node-id>
3. Eliminate Accidental Affinities
Double-check your job config for hidden affinity rules that might pin jobs to node 1. Even a subtle rule like this would cause the clustering you’re seeing:
# Remove this if present! affinity { attribute = "${node.unique.name}" value = "node-1" weight = 100 }
Test the Fix
Once you’ve applied the spread constraint (or adjusted node weights), launch a batch of your count=1 jobs. You should see them distributed evenly across all three nodes instead of clustering on the first one.
内容的提问来源于stack exchange,提问作者imehl

