You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

GCP环境中CDAP搭建BigQuery到GCS的ETL管道问题排查

CDAP Pipeline: BigQuery to GCS Not Generating CSV Files (and Filename Configuration)

Let's break down the issues you ran into with your CDAP pipeline, plus share the fixes you could have used before switching to DataFlow:

1. Why You Got JSON Files Instead of CSV

Looking at your GCS sink configuration, you set "format": "json" — that’s exactly why the output files were in JSON format instead of CSV. For CSV output, you just need to update this property to "format": "csv" (the CDAP Google Cloud plugin v0.12.2 fully supports this format).

2. Where to Set Custom Output Filenames

CDAP doesn’t have a direct "single filename" setting for MapReduce-based pipelines (since MapReduce writes distributed part files by design), but you can control the filename structure with two properties in your GCS sink:

  • Add "filenamePrefix": "your-custom-prefix" to set a base name for your output files.
  • Your existing "suffix": "yyyy-MM-dd" will append the run date to the prefix.

With these, you’d get files like your-custom-prefix-2024-05-20-part-r-00000.csv instead of the generic part files.

3. Corrected GCS Sink Configuration

Here’s the fixed snippet for your GCS sink plugin (key changes are highlighted):

{
  "name": "Google Cloud Storage",
  "plugin": {
    "name": "GCS",
    "type": "batchsink",
    "label": "Google Cloud Storage",
    "artifact": {
      "name": "google-cloud",
      "version": "0.12.2",
      "scope": "SYSTEM"
    },
    "properties": {
      "project": "bi-data-science",
      "suffix": "yyyy-MM-dd",
      "format": "csv", // Changed from json to csv
      "serviceFilePath": "/home/ubuntu/bi-data-science-cdap-4cbf526de374.json",
      "schema": "{\"type\":\"record\",\"name\":\"etlSchemaBody\",\"fields\":[{\"name\":\"destination_name\",\"type\":[\"string\",\"null\"]},{\"name\":\"destination_country\",\"type\":[\"string\",\"null\"]},{\"name\":\"timestamp\",\"type\":[\"double\",\"null\"]},{\"name\":\"desktop\",\"type\":[\"double\",\"null\"]},{\"name\":\"tablet\",\"type\":[\"double\",\"null\"]},{\"name\":\"mobile\",\"type\":[\"double\",\"null\"]}]}",
      "delimiter": ",",
      "referenceName": "gcs_cdap",
      "path": "gs://hurb_sandbox/cdap_experiments/",
      "filenamePrefix": "bq_gcs_export" // Added custom filename prefix
    }
  },
  // Remaining schema/input configuration stays the same
}

4. Bonus: Merging Part Files into a Single CSV

If you needed a single CSV file instead of distributed parts:

  • Switch your pipeline engine to Spark ("engine": "spark" in the top-level config) — Spark can be configured to write a single file for smaller datasets.
  • For MapReduce, you could add a post-action to run gsutil compose to merge the part files, but this requires the CDAP runtime to have access to the gsutil CLI.

Final Note on Switching to DataFlow

It makes total sense to move to DataFlow for this use case — it’s deeply integrated with GCP services like BigQuery and GCS, handles output formatting/naming more intuitively, and removes the overhead of managing an additional CDAP cluster.

内容的提问来源于stack exchange,提问作者David Beyda

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:44:55