请求获取Apache Druid中创建零记录数据源/表的Spec模板
Hey there! Creating a zero-record datasource in Apache Druid is a common need for testing schemas, setting up placeholders for future ingestion, or aligning pipeline structures without actual data. Below is a ready-to-use spec template and a breakdown of each key component to help you adapt it to your needs.
{ "type": "index_parallel", "spec": { "dataSchema": { "dataSource": "zero_record_demo", "timestampSpec": { "column": "event_timestamp", "format": "iso" }, "dimensionsSpec": { "dimensions": [ { "name": "user_id", "type": "string" }, { "name": "product_category", "type": "string" }, { "name": "transaction_amount", "type": "double" } ] }, "metricsSpec": [ { "name": "record_count", "type": "count" }, { "name": "total_amount", "type": "doubleSum", "fieldName": "transaction_amount" } ], "granularitySpec": { "type": "uniform", "segmentGranularity": "day", "queryGranularity": "none", "rollup": false } }, "ioConfig": { "type": "index_parallel", "inputSource": { "type": "inline", "data": "" }, "inputFormat": { "type": "json", "flattenSpec": { "useFieldDiscovery": false } } }, "tuningConfig": { "type": "index_parallel", "partitionsSpec": { "type": "dynamic" } } } }
Let’s break down each critical section so you can tweak this template for your use case:
Task Type (
type): We useindex_parallelhere because it’s the most flexible batch indexing task in Druid. You could also useindex_singleif you prefer, butindex_parallelscales seamlessly if you later add data to this datasource.Data Schema (
dataSchema):dataSource: This is the name of your zero-record table—feel free to rename it to something meaningful for your workflow.timestampSpec: Required for all Druid datasources, even empty ones. Define a timestamp column name and format (we use ISO here, butmillisor other formats work too).dimensionsSpec: List all categorical columns (dimensions) you want your datasource to have, with their respective data types (string, long, double, etc.).metricsSpec: Optional but useful if you want to pre-define aggregation metrics (counts, sums, etc.) that will be ready once you start ingesting data. Even with zero records, these metric definitions persist in the datasource.granularitySpec: Sets segment and query granularity. We disable rollup (rollup: false) since there’s no data to aggregate, but you can adjust this if you plan to use rollup later.
IO Config (
ioConfig):inputSource.type: inline: Uses an inline input source with an empty string fordata—this ensures no records are ingested.inputFormat: Specifies the data format (JSON here) and disables field discovery (useFieldDiscovery: false) to enforce your defined schema instead of letting Druid auto-detect fields (which would fail with empty data).
Tuning Config (
tuningConfig): Dynamic partitioning works perfectly for zero-record tables since there’s no data to split into partitions. You can leave this as-is for most scenarios.
Once you submit this spec via the Druid Console or API, confirm the empty datasource exists by running a simple query:
SELECT * FROM zero_record_demo
This will return an empty result set, confirming your zero-record table is set up correctly.
内容的提问来源于stack exchange,提问作者Moushmi

