You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

咨询Kinesis Data Stream中Shards与Partition Key的概念及通俗解释

通俗理解Kinesis Data Stream的Shards和Partition Key

Hey there! I totally get it—AWS official docs can feel like reading a dense technical textbook sometimes, so let’s ditch the jargon and use a real-world analogy to make this click.

What’s a Shard?

Think of your Kinesis Data Stream as a busy package sorting warehouse. A Shard is like an independent sorting station inside that warehouse. Here’s the breakdown:

  • Each Shard has a fixed "work capacity": it can handle up to 1MB of incoming data (or 1000 individual records) per second, and can serve up to 2MB of data per second to your consumer applications.
  • Shards are completely isolated from each other. If one station is swamped with packages, it won’t slow down the others—you can assign separate teams (consumer apps) to each Shard to split the workload efficiently.
  • Inside a single Shard, data records are stored in strict arrival order. It’s like packages at the station being lined up exactly as they arrived, so your consumer processes them in the same sequence they were sent.

What’s a Partition Key?

Sticking with the warehouse analogy: a Partition Key is like the neighborhood or zip code label on every package you send into the stream. Here’s how it works:

  • When you send a data record to Kinesis, it uses a hash function to turn your Partition Key into a numerical value, then maps that value to a specific Shard.
  • The golden rule: records with the same Partition Key always end up in the same Shard. This is critical if you need to process related records in order—like all actions from a single user (browsing, adding to cart, checking out) will stay grouped together in one Shard, so they never get out of sequence.
  • If you use a wide variety of unique Partition Keys, Kinesis will spread your records evenly across all available Shards. This lets you leverage the full capacity of your stream (since each Shard contributes its own throughput).

How Shards and Partition Keys Work Together

Let’s tie it all together with a concrete example:

  • Suppose you set up a Kinesis Stream with 2 Shards.
  • You’re sending user activity records, each tagged with a user ID as the Partition Key.
  • Kinesis hashes each user ID: IDs starting with A-M get routed to Shard 1, N-Z go to Shard 2.
  • Now you can run two separate consumer apps—one pulling from Shard 1, one from Shard 2—to process twice as much data per second as a single Shard could handle.
  • But if you made the mistake of using the same Partition Key for every record (like "all-users"), every single record would pile into just one Shard. The other Shard would sit idle, and you’d hit the throughput limit of that single Shard way faster than you should.

A few quick extra notes to round it out:

  • You can adjust the number of Shards later (split or merge them) if your data volume grows or shrinks.
  • Each Shard stores data for a set retention period (default 24 hours, max 365 days), so consumers can go back and reprocess old data if needed.

内容的提问来源于stack exchange,提问作者Desp

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 08:02:53