You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Google Dataprep跨数据集复用现有Recipe操作步骤咨询

Reusing a Google Cloud Dataset Recipe Across Datasets

Hey there! I’ve helped a bunch of folks with this exact issue—reusing a dataset recipe across Google Cloud projects/datasets doesn’t have a super obvious step-by-step in the official docs, so let’s break it down clearly. Below are both UI and command-line methods depending on how you prefer to work:

Prerequisites First

Before diving in, make sure you have:

  • Edit access to both the source dataset (where your original recipe lives) and the target dataset.
  • If working cross-project, set up cross-project permissions so the target project’s service account can access any dependencies from the source (like read access to source tables, if needed).

Method 1: Using the Google Cloud Console (UI)

This is great if you prefer clicking through a visual interface:

  • Step 1: Locate your original recipe
    Open the Cloud Console and navigate to the service hosting your recipe (e.g., BigQuery Dataform, Cloud Data Fusion, or BigQuery’s transformation recipes). Find the recipe tied to your source dataset and open its details page.
  • Step 2: Copy/export the recipe configuration
    Look for options like "Export", "Copy", or "Download Config" (labeling varies by service). For code-based recipes (like Dataform’s SQLX files), you can directly copy the code. For pipeline-style recipes (like Cloud Data Fusion), download the full configuration JSON file.
  • Step 3: Create a new recipe for the target dataset
    Navigate to the same service in your target dataset’s project (or same project, different dataset) and click "Create Recipe" or "New Pipeline". Choose the option to "Import from existing config" or paste the copied code directly.
  • Step 4: Update dataset references
    Go through the recipe and replace every reference to the source dataset (e.g., my-source-project.my-source-dataset.my-table) with the target dataset’s full path. Double-check dependencies like output locations, connected data sources, and any environment-specific variables.
  • Step 5: Test and deploy
    Run a test execution of the recipe to make sure it works with the target dataset (no missing permissions or broken references). Once it passes, deploy it to your target dataset’s workflow.

Method 2: Using the gcloud Command Line

Perfect if you want to automate or script the process:

  • Step 1: Export the source recipe
    Use the appropriate gcloud command for your service. Examples:
    • For BigQuery Dataform:
      gcloud dataform repositories export --project=SOURCE_PROJECT --location=SOURCE_REGION --repository=SOURCE_REPO --destination-path=./my-recipe-config
      
    • For Cloud Data Fusion:
      gcloud data-fusion pipelines export --project=SOURCE_PROJECT --location=SOURCE_REGION --instance=SOURCE_INSTANCE --name=SOURCE_PIPELINE --destination=./my-pipeline-recipe.json
      
  • Step 2: Update dataset references
    Open the exported files in a text editor and do a find-and-replace to swap all source dataset paths with the target dataset’s details.
  • Step 3: Import the recipe to the target
    Use the corresponding import command for your service:
    • For BigQuery Dataform:
      gcloud dataform repositories import --project=TARGET_PROJECT --location=TARGET_REGION --repository=TARGET_REPO --source-path=./my-recipe-config
      
    • For Cloud Data Fusion:
      gcloud data-fusion pipelines import --project=TARGET_PROJECT --location=TARGET_REGION --instance=TARGET_INSTANCE --name=TARGET_PIPELINE --source=./my-pipeline-recipe.json
      
  • Step 4: Validate and deploy
    Test the recipe with a dry run or test execution. For Dataform, that might look like:
    gcloud dataform jobs run --project=TARGET_PROJECT --location=TARGET_REGION --repository=TARGET_REPO --compilation-result=latest
    
    If everything checks out, deploy it to production.

Quick Notes to Avoid Headaches

  • Double-check permissions: If cross-project, ensure the target service account has roles/bigquery.dataViewer (or equivalent) on the source dataset if your recipe needs to read from it.
  • Service-specific tweaks: Different GCP tools handle recipes differently—Dataform uses SQLX files, Cloud Data Fusion uses JSON pipelines, so adjust steps to match your tool.
  • Test first: Always run the recipe in a non-production target dataset first to avoid accidental data changes.

内容的提问来源于stack exchange,提问作者BeeKay

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 06:39:21