You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

重启Azure Stream Analytics作业不丢数据:缩放流单元后如何读取未读事件?

Awesome question—this is a common scenario when optimizing Azure Stream Analytics jobs, and luckily the platform has built-in mechanisms to prevent data loss when restarting after scaling stream units. Let’s walk through what you need to know and do:

Key Background Concepts

First, let’s get clear on the core features that make this work:

  • Checkpointing: By default, ASA regularly saves its processing state (including the exact offset it’s read up to for each Event Hub partition) to an Azure Storage account. This is the foundation of resuming without data loss.
  • Event Hub Offsets: Each Event Hub partition tracks events with a unique offset. ASA uses these offsets to pick up exactly where it left off when restarted.
  • Event Retention: Event Hub stores events for a configurable period (default 7 days, up to 90 days for Premium tiers). This ensures events generated while your ASA job is stopped are still available to read when it restarts.
Step-by-Step to Guarantee No Data Loss When Restarting After Scaling

Follow these steps to ensure your job picks up all unprocessed events post-scaling:

  1. Validate Your Event Hub Input Configuration

    • Go to your ASA job’s Inputs blade, select your Event Hub input, and check the Event position setting:
      • For most cases, leave this set to From checkpoint/last processing time (the default). This tells ASA to use the saved checkpoint to resume.
      • If you’re worried about checkpoint corruption (rare, but possible), set it to From earliest event as a fallback—this ensures ASA will read all events in the Event Hub’s retention window if no valid checkpoint exists.
  2. Confirm Checkpointing is Enabled

    • Checkpointing is enabled by default, but it’s worth verifying:
      • Navigate to your ASA job’s Configuration blade, find the Checkpointing section.
      • Ensure it’s set to Enabled, and the linked Azure Storage account is accessible (ASA needs write access to save checkpoints).
  3. Stop the Job Gracefully

    • Always stop your ASA job using the official methods:
      • Azure Portal: Use the Stop button on the job overview page.
      • Azure CLI: Run az stream-analytics job stop --resource-group <your-rg> --name <your-job-name>
    • A graceful stop ensures ASA completes its final checkpoint save, capturing the latest read offsets for all Event Hub partitions.
  4. Adjust Stream Units and Restart

    • After scaling your stream units (up or down), start the job again. ASA will automatically:
      • Load the latest checkpoint data.
      • Resume reading each Event Hub partition from the exact offset it stopped at before the job was shut down.
    • If for some reason the checkpoint is unavailable (e.g., storage account was temporarily down), ASA will fall back to the Event position setting you configured (earliest event, or last processing time).
  5. Verify Processing is Working Correctly

    • After restarting, monitor your job via the Monitoring blade:
      • Check the Input Events Received metric to confirm events are being ingested.
      • Cross-reference with Event Hub’s Incoming Messages metric to ensure all events from the downtime window are being picked up.
    • If you need to manually adjust the starting position (e.g., to reprocess a specific time range), you can temporarily set the Event Hub input’s Event position to Custom offset and specify a timestamp or offset value before restarting.
Critical Additional Tips
  • Don’t Mess With Checkpoint Storage: Avoid deleting the auto-created storage container named azurestreamanalytics-job-<your-job-name>—this is where checkpoints are stored. Deleting it will force ASA to start fresh, potentially causing data loss or duplicate processing.
  • Extend Event Retention if Needed: If you anticipate longer downtime (e.g., more than 7 days), increase your Event Hub’s message retention period to match. This ensures events aren’t expired before ASA can read them.
  • Partition Consistency: ASA processes each Event Hub partition independently, so scaling stream units (which adjusts parallelism) won’t disrupt per-partition offset tracking.

内容的提问来源于stack exchange,提问作者Assasins

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 15:32:43