同一PartitionKey的Kinesis触发Lambda为何并发?如何实现串行处理?
我已将Kinesis Stream配置为AWS Lambda的事件触发器,触发器配置如下:
Activate trigger: Yes Authentication:[{}] Batch size: 10 Batch window: 1 Concurrent batches per shard: 1 Last processing result: OK Maximum age of record: 10800 On-failure destination: None Report batch item failures: No Retry attempts: None Split batch on error: No Starting position: LATEST Tumbling window duration: None UUID: b90da7ec-6e4c-4848-99f5-effa2cf5e0cf
我向该流写入了多条带有同一PartitionKey的记录,数分钟后查看Lambda的Concurrent executions指标,发现存在并发执行的情况。
根据AWS文档说明:
A partition key is used to group data by shard within a stream. Kinesis Data Streams
segregates the data records belonging to a stream into multiple shards. It uses the
partition key that is associated with each data record to determine which shard a
given data record belongs to. Partition keys are Unicode strings, with a maximum
length limit of 256 characters for each key. An MD5 hash function is used to map
partition keys to 128-bit integer values and to map associated data records to
shards using the hash key ranges of the shards. When an application puts data into
a stream, it must specify a partition key.
请问为何我的Lambda会出现并发执行?我该如何配置才能让Lambda对Kinesis分片进行串行无并发的处理?
并发执行的原因
- Kinesis多分片存在:同一PartitionKey的记录只会落在单个分片,但如果你的Kinesis Stream有多个分片(比如自动扩容触发分片分裂、初始创建时就配置多分片),Lambda会为每个分片独立启动执行进程,这是Kinesis与Lambda集成的默认行为,每个分片的处理相互隔离,自然会出现并发。
- Lambda全局并发未限制:即使单分片的并发批次设为1,若Lambda函数未设置全局并发上限,多个分片的执行请求仍会同时触发,导致整体并发。
- 异常执行残留:极少数情况下,若之前的Lambda执行进程未正常终止(比如因异常挂起),可能出现残留并发,但你配置了
Retry attempts: None,这种概率极低。
实现分片串行无并发的配置方法
要实现Lambda对Kinesis分片的全串行处理,需从两方面配置:
1. 确保Kinesis Stream仅含单个分片
- 查看Kinesis控制台的分片数量,若存在多分片,可手动合并分片(仅适用于按需模式或已关闭自动扩容的流);
- 关闭Kinesis Stream的自动分片扩容功能,避免流量增长时自动创建新分片,确保流始终只有一个分片。
2. 配置Lambda限制全局并发
- 保持触发器的
Concurrent batches per shard为1(你当前已正确配置),确保单个分片内的批次串行处理; - 在Lambda控制台设置函数的并发上限为1:进入函数配置的“并发”页面,将“预留并发”设为1,或直接设置函数的“并发上限”为1,强制函数同时只能运行一个执行实例;
- 若无需处理部分失败重试,可保持
Report batch item failures为No,后续若有失败场景,启用该配置可避免不必要的重试并发。
内容的提问来源于stack exchange,提问作者jonua

