如何将yield生成器输出的JSON字符串合并为单个JSON文件?
合并多个JSON对象为结构化JSON文件(Python实现)
我通过yield生成器获取到多个嵌套字典形式的JSON字符串,当前输出为独立的单个JSON对象,希望将它们合并为一个JSON文件。可接受格式为包含这些对象的JSON数组,理想格式是按域名聚合浏览器数据的结构化JSON。
当前输出示例
{"domain.com": {"Chrome": "19362.344607264396"}} {"domain.com": {"ChromeMobile": "7177.498437391487"}} {"another.com": {"MobileSafari": "6237.433155080214"}} {"another.com": {"Safari": "5895.409403430795"}}
可接受合并格式
[ {"domain.com": {"Chrome": "19362.344607264396"}}, {"domain.com": {"ChromeMobile": "7177.498437391487"}}, {"another.com": {"MobileSafari": "6237.433155080214"}}, {"another.com": {"Safari": "5895.409403430795"}} ]
理想合并格式
{ "browsers": [ { "domain.com": { "Chrome": "19362.344607264396", "ChromeMobile": "7177.498437391487" }, "another.com": { "MobileSafari": "6237.433155080214", "Safari": "5895.409403430795" } } ] }
我的Python代码
# Cloudflare zone bandwidth total def browser_map_page_views(domain_zone): cloudflare = prom.custom_query( query="topk(5, sum by(family) (increase(browser_map_page_views_count{job='cloudflare', zone='"f'{domain_zone}'"'}[10d])))" ) for domain_z in cloudflare: user_agent = domain_z['metric']['family'] value = domain_z['value'][1] yield {domain_zone: {user_agent: {'value': value}}} # Get list of zones from Prometheus based on Host Tracker data def domain_zones(): zones_domain = prom.custom_query( query="host_tracker_uptime_percent{job='donodeexporter'}" ) for domain_z in zones_domain: yield domain_z['metric']['zone'] # Get a list of domains and substitution each one into a request of Prometheus query. for domain_list in domain_zones(): for dict in browser_map_page_views(domain_zone=domain_list): dicts = dict print(json.dumps(dicts))
解决方案
注意事项
原代码中browser_map_page_views函数yield的结构是{domain_zone: {user_agent: {'value': value}}},但示例输出里是直接{user_agent: value},如果要和示例格式一致,需要修改yield语句为:
yield {domain_zone: {user_agent: value}}
1. 生成可接受的JSON数组格式
将所有生成的JSON对象收集到列表中,最后统一序列化为JSON数组:
import json # 保留原有的browser_map_page_views和domain_zones函数 # 收集所有对象到列表 result_array = [] for domain in domain_zones(): for item in browser_map_page_views(domain_zone=domain): result_array.append(item) # 打印输出或写入文件 print(json.dumps(result_array, indent=2)) # 写入文件 with open('output_array.json', 'w', encoding='utf-8') as f: json.dump(result_array, f, indent=2, ensure_ascii=False)
2. 生成理想的聚合格式
按域名聚合浏览器数据,构建结构化JSON:
import json # 保留原有的browser_map_page_views和domain_zones函数 # 初始化理想格式的结构 aggregated_result = {"browsers": [{}]} domain_dict = aggregated_result["browsers"][0] for domain in domain_zones(): # 为域名初始化空字典(避免KeyError) if domain not in domain_dict: domain_dict[domain] = {} # 遍历当前域名的所有浏览器数据并合并 for item in browser_map_page_views(domain_zone=domain): domain_dict[domain].update(item[domain]) # 打印输出或写入文件 print(json.dumps(aggregated_result, indent=2)) # 写入文件 with open('output_ideal.json', 'w', encoding='utf-8') as f: json.dump(aggregated_result, f, indent=2, ensure_ascii=False)
内容的提问来源于stack exchange,提问作者Rostyslav Malenko
相关产品推荐
相关产品推荐

