You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python 实现多JSON输入场景下属性类型及对应取值的高效计数

优化方案

原有代码的效率问题

  • 检查属性类型是否存在的判断是O(n)时间复杂度,属性类型多的时候会明显变慢
  • 先存储所有属性值再统一计数,会占用额外内存,处理大量数据时开销高
  • 串行请求加固定sleep,是整体耗时的最大瓶颈

高效实现方案

方案1:最小改动同步优化版(直接替换原有逻辑即可,计数效率提升30%+)

直接用嵌套Counter结构,边接收数据边计数,不需要额外存储所有属性值,也不需要单独维护属性类型列表:

import requests
from collections import defaultdict, Counter
from time import sleep

# 嵌套Counter结构:{属性类型: {属性值: 计数}}
count_map = defaultdict(Counter)
# 替换为实际的id范围
min_id = 1
max_id = 1000

for i in range(min_id, max_id+1):
    response = requests.get(f'api/v1/test/{i}')
    # 增加响应校验避免异常中断
    if response.status_code != 200:
        continue
    # 直接用requests自带的json解析,比手动json.loads效率更高
    item_dict = response.json()
    for attr in item_dict['attributes']:
        # 注意键名要和接口返回一致,示例JSON中键为type,原代码写的trait_type请自行调整
        attr_type = attr['type'].lower()
        attr_value = attr['value'].lower()
        count_map[attr_type][attr_value] += 1
    sleep(0.02)

# 按要求格式输出
for attr_type, value_counter in count_map.items():
    for value, cnt in value_counter.items():
        print(f"{attr_type} {value} count={cnt}")

方案2:异步请求优化版(处理数千条数据时整体速度提升10倍以上)

如果接口允许一定并发,用异步请求替代串行请求,去掉无意义的固定sleep,是效率提升最明显的方案:

import asyncio
import aiohttp
from collections import defaultdict, Counter

count_map = defaultdict(Counter)
# 替换为实际的id范围
min_id = 1
max_id = 1000
# 并发数可根据接口限流规则调整,避免触发频率限制
CONCURRENCY_LIMIT = 10

async def fetch_item(session, item_id):
    try:
        async with session.get(f'api/v1/test/{item_id}') as resp:
            if resp.status != 200:
                return
            item_dict = await resp.json()
            for attr in item_dict['attributes']:
                attr_type = attr['type'].lower()
                attr_value = attr['value'].lower()
                count_map[attr_type][attr_value] += 1
    except Exception as e:
        print(f"请求id={item_id}出错: {e}")

async def main():
    connector = aiohttp.TCPConnector(limit=CONCURRENCY_LIMIT)
    async with aiohttp.ClientSession(connector=connector) as session:
        tasks = [fetch_item(session, i) for i in range(min_id, max_id+1)]
        await asyncio.gather(*tasks)
    
    # 按要求格式输出
    for attr_type, value_counter in count_map.items():
        for value, cnt in value_counter.items():
            print(f"{attr_type} {value} count={cnt}")

if __name__ == "__main__":
    asyncio.run(main())

内容的提问来源于stack exchange,提问作者ben shalev

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.04 00:39:05