You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何按子列表大小拆分字典列表,且保证同邮箱用户在同一子列表?

按子列表大小拆分字典列表,确保同邮箱项在同一子列表

需求说明

给定如下Python字典列表:

total_list = [
    {'email': 'usera@email.com', 'id': 1, 'country': 'UK'},
    {'email': 'usera@email.com', 'id': 1, 'country': 'Germany'}, 
    {'email': 'userb@email.com', 'id': 2, 'country': 'UK'},
    {'email': 'userc@email.com', 'id': 3, 'country': 'Italy'},
    {'email': 'userc@email.com', 'id': 3, 'country': 'Netherland'},
    {'email': 'userd@email.com', 'id': 4, 'country': 'France'},
    # ... 更多元素
]

需要将该列表拆分为多个子列表,要求:

  • 优先保证每个子列表的元素数量不超过指定上限(示例为最多3个)
  • 必须确保所有拥有相同email的字典项处于同一个子列表中

用户曾尝试先按email排序再固定大小拆分的方法,但无法100%满足第二个要求,因此需要可靠的实现方案。

实现方案

核心思路是先按email分组,再将分组后的批次合并到子列表中,确保每个子列表的总元素数不超过上限,同时绝不拆分同一邮箱的分组。

def split_list_by_email_and_size(input_list, max_size=3):
    # 第一步:按email分组,归集同一邮箱的所有项
    email_groups = {}
    for item in input_list:
        email = item['email']
        email_groups.setdefault(email, []).append(item)
    
    # 第二步:合并分组为符合大小要求的子列表
    result = []
    current_batch = []
    current_count = 0
    
    for group in email_groups.values():
        group_size = len(group)
        # 当前批次加该分组会超上限时,先存当前批次,再以该分组为新批次起点
        if current_count + group_size > max_size:
            if current_batch:
                result.append(current_batch)
            current_batch = group
            current_count = group_size
        else:
            # 可以安全加入当前批次
            current_batch.extend(group)
            current_count += group_size
    
    # 处理最后剩余的批次
    if current_batch:
        result.append(current_batch)
    
    return result

# 测试示例
total_list = [
    {'email': 'usera@email.com', 'id': 1, 'country': 'UK'},
    {'email': 'usera@email.com', 'id': 1, 'country': 'Germany'}, 
    {'email': 'userb@email.com', 'id': 2, 'country': 'UK'},
    {'email': 'userc@email.com', 'id': 3, 'country': 'Italy'},
    {'email': 'userc@email.com', 'id': 3, 'country': 'Netherland'},
    {'email': 'userd@email.com', 'id': 4, 'country': 'France'},
]

# 按每个子列表最多3个元素拆分
split_result = split_list_by_email_and_size(total_list, max_size=3)
for idx, batch in enumerate(split_result, 1):
    print(f"子列表 {idx}: {batch}")

逻辑说明

  1. 分组阶段:遍历原列表,把相同email的项归类到同一列表,从根源避免同邮箱项被拆分。
  2. 合并阶段:依次将每个分组加入当前子列表,若加入后超过max_size,则先保存当前子列表,再将该分组作为新的子列表起点;否则直接合并到当前子列表。
  3. 特殊情况处理:如果某个邮箱对应的分组大小超过max_size(比如某邮箱有4个项),该分组会单独成为一个子列表——这是符合需求优先级的,因为必须保证同邮箱项不拆分。

测试输出

运行上述代码后,输出结果如下:

子列表 1: [{'email': 'usera@email.com', 'id': 1, 'country': 'UK'}, {'email': 'usera@email.com', 'id': 1, 'country': 'Germany'}, {'email': 'userb@email.com', 'id': 2, 'country': 'UK'}]
子列表 2: [{'email': 'userc@email.com', 'id': 3, 'country': 'Italy'}, {'email': 'userc@email.com', 'id': 3, 'country': 'Netherland'}, {'email': 'userd@email.com', 'id': 4, 'country': 'France'}]

所有同邮箱的项都在同一个子列表中,且每个子列表的元素数量未超过指定上限。

内容的提问来源于stack exchange,提问作者Flora Biletsiou

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 16:20:41