You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中合并ThreadPoolExecutor返回的JSON对象?

合并ThreadPoolExecutor返回的字典结果

问题背景

需要合并ThreadPoolExecutor返回的results中的所有字典元素,results的具体数据如下:

  • element 0:
    {'00000001': {'foo_id': 83959370, 'bar_id': 'ABCD1'}, '00000002': {'foo_id': 83959371, 'bar_id': 'ABCD2'}}
    
  • element 1:
    {'00000003': {'foo_id': 83959372, 'bar_id': 'ABCD2'}, '00000004': {'foo_id': 83959373, 'bar_id': 'ABCD4'}}
    
  • element 2:
    {'00000005': {'foo_id': 83959374, 'bar_id': 'ABCD3'}, '00000006': {'foo_id': 83959375, 'bar_id': 'ABCD6'}}
    

当前使用ThreadPoolExecutor的代码:

from concurrent.futures import ThreadPoolExecutor
import logging

logger = logging.getLogger(__name__)

# 假设send_request和foo_ids_chunks、authorisation_token已定义
with ThreadPoolExecutor(max_workers=len(foo_ids_chunks)) as executor:
        # 并行调用send_request处理每个chunk
        results = executor.map(send_request, foo_ids_chunks, [authorisation_token] * len(foo_ids_chunks))
        logger.info(f"results {results}")
        # 遍历结果打日志
        for r in results:
            logger.info(f"result {r}")

尝试过两种无效的合并方案:

  1. 错误的列表推导式(逻辑不符,结果是字典而非列表):
    merged_results = [item for sublist in results for item in sublist]
    
  2. 字典推导式(因迭代器耗尽导致无数据):
    merged_results = {key: value for json_dict in results for key, value in json_dict.items()}
    

期望得到的合并结果:

{
  "00000001":{
    "foo_id":83959370,
    "bar_id":"ABCD1"
  },
  "00000002":{
    "foo_id":83959371,
    "bar_id":"ABCD2"
  },
  "00000003":{
    "foo_id":83959372,
    "bar_id":"ABCD2"
  },
  "00000004":{
    "foo_id":83959373,
    "bar_id":"ABCD4"
  },
  "00000005":{
    "foo_id":83959374,
    "bar_id":"ABCD3"
  },
  "00000006":{
    "foo_id":83959375,
    "bar_id":"ABCD6"
  }
}

问题原因

executor.map()返回的是迭代器而非列表,迭代器只能被遍历一次:代码中已经通过for r in results遍历打日志,后续再用推导式遍历results时,迭代器已耗尽,无法获取数据。此外第一个列表推导式逻辑错误,遍历字典默认只会拿到键,不是目标键值对。

正确实现

先将迭代器转换为列表保存,再进行合并操作:

with ThreadPoolExecutor(max_workers=len(foo_ids_chunks)) as executor:
        results = executor.map(send_request, foo_ids_chunks, [authorisation_token] * len(foo_ids_chunks))
        # 转换为列表,避免迭代器耗尽
        results_list = list(results)
        logger.info(f"results {results_list}")
        
        # 遍历列表打日志
        for r in results_list:
            logger.info(f"result {r}")
        
        # 方案1:用update方法合并字典
        merged_results = {}
        for json_dict in results_list:
            merged_results.update(json_dict)
        
        # 方案2:用字典推导式(基于保存的列表)
        # merged_results = {k: v for d in results_list for k, v in d.items()}

结果验证

运行后merged_results会包含所有子字典的键值对,与期望结构完全一致。若存在重复键,后续字典的键值对会覆盖前面的(示例数据无重复键,无需额外处理)。

内容的提问来源于stack exchange,提问作者Stéphane GRILLON

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.09 04:51:12