集成Tika Parser后StormCrawler性能异常问题排查求助
Based on your description and topology config, here’s a breakdown of the likely bottlenecks and actionable fixes:
1. Immediate Fix: Increase Tika Parser Bolt Parallelism
Your current topology sets parser_bolt (Tika) to parallelism: 1, while JSoup’s parse bolt has 5 instances. This is a critical bottleneck—all non-HTML documents (PDFs, docs) have to go through a single Tika bolt, which can’t keep up with the volume from the fetcher and JSoup parser.
Adjust the parallelism to match or exceed JSoup’s count (start with 5, then scale based on your CPU cores):
- id: "parser_bolt" className: "com.digitalpebble.stormcrawler.tika.ParserBolt" parallelism: 5 # Increase this to distribute the workload
This spreads the resource-heavy Tika parsing across multiple threads, preventing backpressure from building up and crashing overall performance.
2. Fix Memory Leaks & Reduce GC Pressure
Your OOM errors (even with 4GB heap) and memory usage spiking to 10-15GB suggest a memory leak in the Tika integration. Here’s how to address it:
- Reuse Tika Parser Instances: Ensure the Tika bolt uses a parser pool instead of creating new parsers for every document. Add this to your
crawler-custom-conf.yaml:
This cuts down on object creation and garbage collection (GC) overhead.tika.parser.pool.size: 10 # Match or exceed your bolt parallelism - Tune JVM Memory: Instead of limiting memory too tightly, set
workers.heap.sizeto 8-12GB (since your workers were already using 10-15GB) and use a GC optimized for large heaps:
This reduces GC thrashing (the cause of high CPU when memory is constrained).worker.childopts: "-XX:+UseG1GC -XX:MaxGCPauseMillis=200 -Xmx8g" - Verify Resource Cleanup: Tika can hold onto InputStreams or parsing contexts if not closed correctly. Double-check that the
ParserBoltis closing all resources after parsing (StormCrawler’s default implementation should handle this, but confirm if you’ve customized it).
3. Validate Stream Routing to Tika
Ensure the RedirectionBolt is only sending non-HTML documents to the Tika bolt. If HTML docs are accidentally routed to Tika, it adds unnecessary load. Confirm the bolt filters based on Content-Type headers (e.g., application/pdf, application/vnd.ms-word) before emitting to the "tika" stream.
4. Identify Slow Parsing Cases
High CPU usage without errors often points to individual documents taking too long to parse. Add logging for parse duration in the Tika bolt to find outliers (large PDFs, docs with embedded media, or corrupted files). You can then:
- Set a timeout for Tika parsing to prevent long-running tasks from blocking the bolt:
tika.parser.timeout: 30000 # 30 seconds - Exclude particularly large or problematic file types if they’re not critical to your use case.
5. Monitor Storm Topology Metrics
Use Storm UI to diagnose backpressure:
- If the queue before
parser_boltis consistently growing, that confirms it’s the bottleneck (further increasing parallelism will help). - Check the
fetcherbolt’s throughput—if it drops when Tika is active, backpressure from the parser bolt is slowing down the entire topology.
内容的提问来源于stack exchange,提问作者elgato

