You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在DynamoDB分页Scan中使用StartingToken及相关问题

问题解答

1. 获取分页所需的NextToken

DynamoDB分页器返回的page对象里的LastEvaluatedKey就是下一页的起始标记,但不能直接传给StartingToken——因为StartingToken需要的是经过分页器编码后的字符串,而非原始的字典结构。

正确的做法是用分页器的encode_token()方法对LastEvaluatedKey进行编码,得到的字符串就是可以返回给客户端的NextToken。示例代码:

import boto3
import json

table = "TableName"
client = boto3.client("dynamodb")
paginator = client.get_paginator("query")
token = None
size = 1

# 初始化测试数据(可忽略)
for i in range(1, 10):
    client.put_item(TableName=table, Item={"PK": {"S": str(i)}, "SK": {"S": str(i)}})

it = paginator.paginate(
    TableName=table,
    ProjectionExpression="PK,SK",
    PaginationConfig={"MaxItems": 100, "PageSize": size, "StartingToken": token}
)

# 获取第一页数据
page = next(it)
print(json.dumps(page, indent=2))

# 生成可供客户端使用的NextToken(如果存在下一页)
if "LastEvaluatedKey" in page:
    next_token = paginator.encode_token(page["LastEvaluatedKey"])
    print("NextToken:", next_token)

客户端下次请求时,直接将这个next_token作为StartingToken传入即可。

2. 不用循环直接获取单页数据

你之前尝试next(it)无效,大概率是因为代码里先写了for page in it再break,导致迭代器已经被推进过一次。直接用next()获取单页的正确方式是:不要提前遍历迭代器,直接调用next()拿到第一页,示例见上面的代码。

如果需要处理“没有更多数据”的情况,可以捕获StopIteration异常:

try:
    page = next(it)
except StopIteration:
    # 无更多数据时的处理逻辑
    page = None

3. LastEvaluatedKey报错的解决

直接把LastEvaluatedKey传给StartingToken会报错,因为分页器的StartingToken要求的是编码后的字符串,而非原始的键字典。必须通过paginator.encode_token()处理后才能正常使用,这也是第一点提到的核心解决方案。

4. 扫描时找到指定数量结果就停止

如果用普通scan而非分页器,可以手动控制扫描流程:每次调用scan获取一页数据,过滤出符合条件的条目,累计数量达到10条就停止,同时记录LastEvaluatedKey以便后续继续扫描。示例代码:

import boto3

client = boto3.client("dynamodb")
table = "TableName"
target_count = 10  # 要找的目标条目数
max_scan_limit = 1000  # 最多扫描的条目数
scanned_total = 0
found_items = []
last_key = None

while len(found_items) < target_count and scanned_total < max_scan_limit:
    scan_params = {
        "TableName": table,
        "FilterExpression": "你的过滤条件",  # 替换为实际过滤规则
        "ProjectionExpression": "PK,SK"
    }
    if last_key:
        scan_params["ExclusiveStartKey"] = last_key
    
    response = client.scan(**scan_params)
    # 若FilterExpression已完成过滤,直接取Items即可;否则可在这里加额外过滤逻辑
    filtered = [item for item in response["Items"] if 你的自定义过滤规则]
    
    found_items.extend(filtered)
    scanned_total += response["ScannedCount"]
    last_key = response.get("LastEvaluatedKey")
    
    # 截断到目标数量
    if len(found_items) > target_count:
        found_items = found_items[:target_count]

# 输出结果和下次扫描的起始键(如果需要)
print("找到的条目:", found_items)
next_scan_key = last_key if len(found_items) == target_count and scanned_total < max_scan_limit else None

这种方式可以在找到10条符合条件的数据后立即停止扫描,无需扫完1000条上限。

内容的提问来源于stack exchange,提问作者rfg

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 23:10:26