如何在Python中合并ThreadPoolExecutor返回的JSON对象?
合并ThreadPoolExecutor返回的字典结果
问题背景
需要合并ThreadPoolExecutor返回的results中的所有字典元素,results的具体数据如下:
- element 0:
{'00000001': {'foo_id': 83959370, 'bar_id': 'ABCD1'}, '00000002': {'foo_id': 83959371, 'bar_id': 'ABCD2'}} - element 1:
{'00000003': {'foo_id': 83959372, 'bar_id': 'ABCD2'}, '00000004': {'foo_id': 83959373, 'bar_id': 'ABCD4'}} - element 2:
{'00000005': {'foo_id': 83959374, 'bar_id': 'ABCD3'}, '00000006': {'foo_id': 83959375, 'bar_id': 'ABCD6'}}
当前使用ThreadPoolExecutor的代码:
from concurrent.futures import ThreadPoolExecutor import logging logger = logging.getLogger(__name__) # 假设send_request和foo_ids_chunks、authorisation_token已定义 with ThreadPoolExecutor(max_workers=len(foo_ids_chunks)) as executor: # 并行调用send_request处理每个chunk results = executor.map(send_request, foo_ids_chunks, [authorisation_token] * len(foo_ids_chunks)) logger.info(f"results {results}") # 遍历结果打日志 for r in results: logger.info(f"result {r}")
尝试过两种无效的合并方案:
- 错误的列表推导式(逻辑不符,结果是字典而非列表):
merged_results = [item for sublist in results for item in sublist] - 字典推导式(因迭代器耗尽导致无数据):
merged_results = {key: value for json_dict in results for key, value in json_dict.items()}
期望得到的合并结果:
{ "00000001":{ "foo_id":83959370, "bar_id":"ABCD1" }, "00000002":{ "foo_id":83959371, "bar_id":"ABCD2" }, "00000003":{ "foo_id":83959372, "bar_id":"ABCD2" }, "00000004":{ "foo_id":83959373, "bar_id":"ABCD4" }, "00000005":{ "foo_id":83959374, "bar_id":"ABCD3" }, "00000006":{ "foo_id":83959375, "bar_id":"ABCD6" } }
问题原因
executor.map()返回的是迭代器而非列表,迭代器只能被遍历一次:代码中已经通过for r in results遍历打日志,后续再用推导式遍历results时,迭代器已耗尽,无法获取数据。此外第一个列表推导式逻辑错误,遍历字典默认只会拿到键,不是目标键值对。
正确实现
先将迭代器转换为列表保存,再进行合并操作:
with ThreadPoolExecutor(max_workers=len(foo_ids_chunks)) as executor: results = executor.map(send_request, foo_ids_chunks, [authorisation_token] * len(foo_ids_chunks)) # 转换为列表,避免迭代器耗尽 results_list = list(results) logger.info(f"results {results_list}") # 遍历列表打日志 for r in results_list: logger.info(f"result {r}") # 方案1:用update方法合并字典 merged_results = {} for json_dict in results_list: merged_results.update(json_dict) # 方案2:用字典推导式(基于保存的列表) # merged_results = {k: v for d in results_list for k, v in d.items()}
结果验证
运行后merged_results会包含所有子字典的键值对,与期望结构完全一致。若存在重复键,后续字典的键值对会覆盖前面的(示例数据无重复键,无需额外处理)。
内容的提问来源于stack exchange,提问作者Stéphane GRILLON
相关产品推荐
相关产品推荐

