关于Druid集群Kafka摄入任务失败及数据完整性的技术咨询
Troubleshooting Druid Indexing Service Task Failures & Data Integrity Checks
First off, let’s tackle the task failure issue since that’s the root of your data integrity concerns. Even without fatal errors in logs, there are several common culprits to look for:
1. Diagnosing Task Failures (No Fatal Log Errors)
- Resource Starvation: MiddleManagers often fail tasks silently if they don’t have enough memory or CPU. Check your
druid.worker.capacity(number of concurrent tasks per MiddleManager) anddruid.indexer.runner.javaOpts(heap allocation) settings. Look for warnings in MiddleManager logs about GC pauses, memory pressure, or task timeouts—these are easy to miss but often the cause. - Kafka Consumer Glitches: Tasks might be hitting Kafka rebalances, offset timeouts, or missing offset ranges. Check if your Kafka topic has high lag (use
kafka-consumer-groups.shto verify). Also, review your ingestion spec’sconsumerProperties—settings likesession.timeout.msormax.poll.interval.msthat are too low can trigger unexpected task failures. - Transient Storage/Metadata Issues: Partial segment writes suggest intermittent problems with Deep Storage (S3/HDFS) or Metadata Store (MySQL/Postgres). Look for logs about failed write attempts or connection timeouts to these services—even non-fatal retries can eventually cause tasks to fail.
- Retry Limits Exceeded: Druid’s default retry count might be too low for your environment. Check
druid.indexer.runner.maxRetriesin the Overlord config. If tasks are failing due to transient issues, bumping this up temporarily can help, but make sure you address the underlying cause first. - Quiet Data Parsing Errors: Sometimes invalid records don’t throw fatal errors but cause partial failures. Enable debug logging for the ingestion task (set
log4j.logger.org.apache.druid.data.input=DEBUG) to catch warnings about parse failures or dropped rows.
2. Verifying All Kafka Data Is Ingested into Segments
To confirm you haven’t lost data, use these methods:
- Offset Range Comparison:
- For each ingestion task, pull the start/end Kafka offsets from task logs (look for lines like "Ingesting from offset X to Y").
- Use Kafka’s
GetOffsetShelltool to map Druid segment timestamps back to Kafka offsets:kafka-run-class.sh kafka.tools.GetOffsetShell --broker-list <kafka-broker>:9092 --topic <your-topic> --time <segment-start-timestamp-ms> --offsets 1 - Cross-check that the total offset range covered by all successful segments matches the total offsets consumed from Kafka.
- Record Count Validation:
- Calculate total Kafka records for your target time range using the offset shell (subtract start offsets from end offsets).
- Run a count query in Druid:
A small gap is normal if you’re filtering records, but a large discrepancy means data is missing.SELECT COUNT(*) FROM <your-datasource> WHERE __time BETWEEN TIMESTAMP '2024-01-01T00:00:00Z' AND TIMESTAMP '2024-01-02T00:00:00Z'
- Task Report Aggregation:
- In the Druid Console, go to the Tasks page and pull the "Task Report" for each failed/successful task. Look for metrics like
ingestedRows,droppedRows, andparseErrors. - Sum these metrics across all tasks to see how much data was processed vs. how much should have been ingested.
- In the Druid Console, go to the Tasks page and pull the "Task Report" for each failed/successful task. Look for metrics like
3. Next Steps to Dig Deeper
- Crank Up Logging: Temporarily set
log4j.logger.org.apache.druid.indexing=DEBUGin your MiddleManager’s log4j config. This will capture granular details about task execution that might be hidden at the default info level. - Check Task Patterns: Use the Overlord API to list all tasks and their statuses:
Look for patterns—do failures happen on specific MiddleManagers, or during peak traffic?curl -X GET http://<overlord-host>:8090/druid/indexer/v1/tasks - Validate Your Ingestion Spec: Double-check your
tuningConfigsettings. For example, ifmaxRowsPerSegmentis too high, tasks might time out before completing. EnsuretaskTimeoutis set appropriately for your data volume.
内容的提问来源于stack exchange,提问作者Frank
相关产品推荐
相关产品推荐

