Google Dataprep跨数据集复用现有Recipe操作步骤咨询
Reusing a Google Cloud Dataset Recipe Across Datasets
Hey there! I’ve helped a bunch of folks with this exact issue—reusing a dataset recipe across Google Cloud projects/datasets doesn’t have a super obvious step-by-step in the official docs, so let’s break it down clearly. Below are both UI and command-line methods depending on how you prefer to work:
Prerequisites First
Before diving in, make sure you have:
- Edit access to both the source dataset (where your original recipe lives) and the target dataset.
- If working cross-project, set up cross-project permissions so the target project’s service account can access any dependencies from the source (like read access to source tables, if needed).
Method 1: Using the Google Cloud Console (UI)
This is great if you prefer clicking through a visual interface:
- Step 1: Locate your original recipe
Open the Cloud Console and navigate to the service hosting your recipe (e.g., BigQuery Dataform, Cloud Data Fusion, or BigQuery’s transformation recipes). Find the recipe tied to your source dataset and open its details page. - Step 2: Copy/export the recipe configuration
Look for options like "Export", "Copy", or "Download Config" (labeling varies by service). For code-based recipes (like Dataform’s SQLX files), you can directly copy the code. For pipeline-style recipes (like Cloud Data Fusion), download the full configuration JSON file. - Step 3: Create a new recipe for the target dataset
Navigate to the same service in your target dataset’s project (or same project, different dataset) and click "Create Recipe" or "New Pipeline". Choose the option to "Import from existing config" or paste the copied code directly. - Step 4: Update dataset references
Go through the recipe and replace every reference to the source dataset (e.g.,my-source-project.my-source-dataset.my-table) with the target dataset’s full path. Double-check dependencies like output locations, connected data sources, and any environment-specific variables. - Step 5: Test and deploy
Run a test execution of the recipe to make sure it works with the target dataset (no missing permissions or broken references). Once it passes, deploy it to your target dataset’s workflow.
Method 2: Using the gcloud Command Line
Perfect if you want to automate or script the process:
- Step 1: Export the source recipe
Use the appropriategcloudcommand for your service. Examples:- For BigQuery Dataform:
gcloud dataform repositories export --project=SOURCE_PROJECT --location=SOURCE_REGION --repository=SOURCE_REPO --destination-path=./my-recipe-config - For Cloud Data Fusion:
gcloud data-fusion pipelines export --project=SOURCE_PROJECT --location=SOURCE_REGION --instance=SOURCE_INSTANCE --name=SOURCE_PIPELINE --destination=./my-pipeline-recipe.json
- For BigQuery Dataform:
- Step 2: Update dataset references
Open the exported files in a text editor and do a find-and-replace to swap all source dataset paths with the target dataset’s details. - Step 3: Import the recipe to the target
Use the corresponding import command for your service:- For BigQuery Dataform:
gcloud dataform repositories import --project=TARGET_PROJECT --location=TARGET_REGION --repository=TARGET_REPO --source-path=./my-recipe-config - For Cloud Data Fusion:
gcloud data-fusion pipelines import --project=TARGET_PROJECT --location=TARGET_REGION --instance=TARGET_INSTANCE --name=TARGET_PIPELINE --source=./my-pipeline-recipe.json
- For BigQuery Dataform:
- Step 4: Validate and deploy
Test the recipe with a dry run or test execution. For Dataform, that might look like:
If everything checks out, deploy it to production.gcloud dataform jobs run --project=TARGET_PROJECT --location=TARGET_REGION --repository=TARGET_REPO --compilation-result=latest
Quick Notes to Avoid Headaches
- Double-check permissions: If cross-project, ensure the target service account has
roles/bigquery.dataViewer(or equivalent) on the source dataset if your recipe needs to read from it. - Service-specific tweaks: Different GCP tools handle recipes differently—Dataform uses SQLX files, Cloud Data Fusion uses JSON pipelines, so adjust steps to match your tool.
- Test first: Always run the recipe in a non-production target dataset first to avoid accidental data changes.
内容的提问来源于stack exchange,提问作者BeeKay
相关产品推荐
相关产品推荐

