You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

请求获取Apache Druid中创建零记录数据源/表的Spec模板

Hey there! Creating a zero-record datasource in Apache Druid is a common need for testing schemas, setting up placeholders for future ingestion, or aligning pipeline structures without actual data. Below is a ready-to-use spec template and a breakdown of each key component to help you adapt it to your needs.

Zero-Record Datasource Spec Template
{
  "type": "index_parallel",
  "spec": {
    "dataSchema": {
      "dataSource": "zero_record_demo",
      "timestampSpec": {
        "column": "event_timestamp",
        "format": "iso"
      },
      "dimensionsSpec": {
        "dimensions": [
          {
            "name": "user_id",
            "type": "string"
          },
          {
            "name": "product_category",
            "type": "string"
          },
          {
            "name": "transaction_amount",
            "type": "double"
          }
        ]
      },
      "metricsSpec": [
        {
          "name": "record_count",
          "type": "count"
        },
        {
          "name": "total_amount",
          "type": "doubleSum",
          "fieldName": "transaction_amount"
        }
      ],
      "granularitySpec": {
        "type": "uniform",
        "segmentGranularity": "day",
        "queryGranularity": "none",
        "rollup": false
      }
    },
    "ioConfig": {
      "type": "index_parallel",
      "inputSource": {
        "type": "inline",
        "data": ""
      },
      "inputFormat": {
        "type": "json",
        "flattenSpec": {
          "useFieldDiscovery": false
        }
      }
    },
    "tuningConfig": {
      "type": "index_parallel",
      "partitionsSpec": {
        "type": "dynamic"
      }
    }
  }
}
Key Spec Components Explained

Let’s break down each critical section so you can tweak this template for your use case:

  • Task Type (type): We use index_parallel here because it’s the most flexible batch indexing task in Druid. You could also use index_single if you prefer, but index_parallel scales seamlessly if you later add data to this datasource.

  • Data Schema (dataSchema):

    • dataSource: This is the name of your zero-record table—feel free to rename it to something meaningful for your workflow.
    • timestampSpec: Required for all Druid datasources, even empty ones. Define a timestamp column name and format (we use ISO here, but millis or other formats work too).
    • dimensionsSpec: List all categorical columns (dimensions) you want your datasource to have, with their respective data types (string, long, double, etc.).
    • metricsSpec: Optional but useful if you want to pre-define aggregation metrics (counts, sums, etc.) that will be ready once you start ingesting data. Even with zero records, these metric definitions persist in the datasource.
    • granularitySpec: Sets segment and query granularity. We disable rollup (rollup: false) since there’s no data to aggregate, but you can adjust this if you plan to use rollup later.
  • IO Config (ioConfig):

    • inputSource.type: inline: Uses an inline input source with an empty string for data—this ensures no records are ingested.
    • inputFormat: Specifies the data format (JSON here) and disables field discovery (useFieldDiscovery: false) to enforce your defined schema instead of letting Druid auto-detect fields (which would fail with empty data).
  • Tuning Config (tuningConfig): Dynamic partitioning works perfectly for zero-record tables since there’s no data to split into partitions. You can leave this as-is for most scenarios.

Quick Verification

Once you submit this spec via the Druid Console or API, confirm the empty datasource exists by running a simple query:

SELECT * FROM zero_record_demo

This will return an empty result set, confirming your zero-record table is set up correctly.

内容的提问来源于stack exchange,提问作者Moushmi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:41:04