Storm Topology设置Bolt B并行度为1200时无法启动问题
Let’s break down what’s going on here and walk through actionable fixes to get your topology up and running with the higher parallelism.
Environment Context
First, let’s recap your setup to make sure we’re on the same page:
- Storm version dependencies:
"org.apache.storm" % "storm-core" % "1.2.1" "org.apache.storm" % "storm-kafka-client" % "1.2.1" - Topology structure: 3 Bolts (A, B, C) where Bolt B is the heavy hitter (~450ms avg processing time; A/C are <1ms)
- Working parallelism: A:4, B:700, C:10
- Failing parallelism: A:4, B:1200, C:10 (no other config changes)
- Key configs:
topology.worker.childopts: -Xms1g -Xmx16g topology.worker.logwriter.childopts: -Xmx1024
What You’re Seeing
When you crank Bolt B to 1200 parallelism:
- The topology keeps trying to load Bolt B executors (you can see this in the repeated "Loading executor B:[XXX XXX]" logs)
- Worker processes restart nonstop, but there’s no obvious error in the main topology or Storm logs
- The critical clue: Your worker exits with code 143 (from the supervisor logs)
Why This Is Happening
Exit code 143 isn’t random—it means the worker process got terminated by a SIGTERM signal (128 + 15 = 143). In Storm, this almost always boils down to two scenarios:
- Supervisor thinks the worker is unresponsive: Loading 1200 executors takes way longer than Storm’s default startup timeout. The supervisor doesn’t get a heartbeat in time, so it kills the worker and tries again.
- Resource exhaustion: Scaling to 1200 executors is pushing past your cluster’s (or individual worker’s) resource limits—think too many threads, file handles, or memory pressure during initialization. The OS or supervisor pulls the plug to prevent instability.
Fixes to Try (Start with the Simplest First)
1. Give Workers More Time to Initialize
Storm’s default timeouts are designed for typical setups, not loading 1200 executors. Extend these settings to give your workers enough time to spin up:
Add these to your topology configuration:
# Give workers 3 minutes to start up (default is 60s) supervisor.worker.start.timeout.secs: 180 # Extend heartbeat timeout to 1 minute (default is 30s) supervisor.worker.timeout.secs: 60 # Make sure message timeout is generous enough (prevents premature kills during startup) topology.message.timeout.secs: 300
2. Spread Executors Across More Workers
1200 executors crammed into too few workers is a recipe for startup overload. Use the topology.workers config to split the load:
For example, if you set topology.workers: 20, each worker will handle ~60 Bolt B executors instead of hundreds. This reduces the initialization strain on any single worker process.
3. Check and Increase OS Resource Limits
Storm needs plenty of file descriptors and threads when running hundreds of executors. Here’s what to do:
- File descriptors: Run
ulimit -nas the Storm user. If it’s under 65535, increase it in/etc/security/limits.conf(add lines likestorm soft nofile 65535andstorm hard nofile 65535) - Thread limits: Ensure the Storm user has enough threads allocated. You can check this with
ulimit -uand adjust it in the same limits.conf file if needed.
4. Add GC Logging to Catch Memory Issues
Even though your Xmx16g is large, initialization of 1200 executors might be causing GC thrashing or out-of-memory errors that aren’t showing up in standard logs. Add GC logging to your worker opts:
topology.worker.childopts: -Xms1g -Xmx16g -XX:+PrintGCDetails -XX:+PrintGCTimeStamps -Xloggc:/var/log/storm/gc-worker.log
Check the GC logs after a failed startup—if you see frequent full GCs or out-of-memory errors, you might need to tweak heap settings or split executors further.
5. Enable Verbose Logging for Hidden Errors
Sometimes the real error is hidden in debug logs. Update your log4j2.xml to set the supervisor log level to DEBUG:
<Logger name="org.apache.storm.daemon.supervisor" level="DEBUG"/>
Also, check the worker’s stderr logs (usually in storm/logs/workers-artifacts/<your-topology-id>/<port>/worker.log)—these often contain JVM-level errors that don’t make it to the main Storm logs.
Final Thoughts
Since your topology works perfectly at 700 parallelism, the issue is definitely tied to the increased startup load of 1200 Bolt B executors. Start with extending the timeouts (that’s the quickest win) and then move to splitting executors across more workers if needed. Don’t overlook OS resource limits—those are a common culprit for silent worker kills.
内容的提问来源于stack exchange,提问作者Saurabh

