You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Google Dataprep:更新数据源调度与GCS触发流及每日取数咨询

Awesome questions—let’s walk through exactly how to handle both scenarios with Google Dataprep, since you’re totally right that running jobs on stale data defeats the purpose of scheduling!

1. Scheduling Jobs Based on Updated Data Sources

Dataprep gives you two solid approaches to ensure your jobs only run against fresh data:

  • Time-based scheduling with dynamic source filtering: If you prefer regular runs (like daily), point your dataset to a GCS directory (not a single fixed file). Then, add a custom filter in your dataset’s import settings to only pull files modified after your last job run. Use Dataprep’s built-in lastModified() function here—for example, lastModified() >= addDays(now(), -1) for daily runs, which grabs only files updated in the last 24 hours. This way, even if you schedule daily runs, you won’t reprocess old data.
  • Event-driven triggering (via Cloud Functions): For true "data updates → job runs" automation, pair GCS with Cloud Functions. When a new file finishes uploading to your target bucket (triggered by the OBJECT_FINALIZE event), the Cloud Function calls Dataprep’s jobs.create API to kick off your job immediately. Just make sure the Cloud Function’s service account has the right Dataprep permissions (like Dataprep Job Runner) to submit jobs.

2. Triggering Dataprep on GCS File Uploads or Daily Runs with Latest Files

Let’s break down both options you asked about:

  • Option 1: Trigger on GCS file uploads
    This is the event-driven flow we mentioned earlier. Here’s a quick breakdown of the steps:
    1. In the Google Cloud Console, create a Cloud Function with a GCS trigger. Select your target bucket and set the event type to Object finalized (when a file is fully uploaded).
    2. Write a simple function (Python or Node.js works) that calls Dataprep’s REST API to start your pre-configured job. You’ll need to authenticate the API call using the Cloud Function’s default service account.
    3. Test by uploading a file to your GCS bucket—your Dataprep job should kick off automatically.
  • Option 2: Daily runs that only process new files
    If you’d rather stick to a daily schedule but avoid reprocessing old data:
    1. Set up your Dataprep dataset to point to your GCS directory (not individual files).
    2. Add a filter rule in the dataset’s import settings to target only recent files. Use lastModified() to filter for files changed since the last run—adjust the time window to match your schedule (e.g., lastModified() >= addHours(now(), -24) for daily runs).
    3. Configure your job to run daily in Dataprep’s scheduler. Make sure to enable any "refresh source data" options if available, so Dataprep rescans the directory each time it runs.
      Bonus: If your data has a built-in timestamp column (like a created_at field), use Dataprep’s Incremental Load feature to automatically pull only new rows added since the last job execution.

Both approaches ensure you’re only processing fresh data, so you won’t waste resources running jobs that produce identical results.

内容的提问来源于stack exchange,提问作者stkvtflw

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 04:34:09