You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何优化10MB JSON数据的条件提取与排序效率?

优化方案:从10MB JSON接口高效筛选目标数据

原代码的核心问题

  1. 全量拉取冗余数据:每次下载10MB JSON,大部分数据不符合筛选条件,浪费带宽与解析时间
  2. 语法与逻辑错误:
    • 异常捕获except data == []完全错误,这不是合法的异常类型,应该是判断data是否为空后切换备用接口
    • 排序lambda中错误引用全局data,而非当前条目sorted_data,会导致解析失败
    • 仅遍历排序后的前20条,可能遗漏符合条件的当日数据
  3. 低效的Checkpoint操作:每次加载都读取整个文本文件,每次写入都打开文件追加,频繁IO拖慢速度
  4. 不必要的全量排序:对所有数据排序,而不是仅筛选符合条件的条目后排序

针对性优化措施

1. 减少网络传输:增量判断+流式解析

  • 用HTTP缓存头避免重复下载:通过If-Modified-Since或ETag头,询问服务器数据是否更新,未更新则直接复用本地缓存
  • 流式JSON解析:使用ijson库边下载边解析,无需把10MB数据全加载到内存,遇到符合条件的条目立即处理

2. 优化本地筛选逻辑

  • 预过滤当日数据:提前计算当日的时间范围,解析creationDate时直接判断是否属于当日,跳过非目标数据
  • 延迟排序:仅对符合status == 6且为当日的条目排序,而非全量排序
  • 内存缓存Checkpoint:把已处理的job_id存在内存集合中,定期持久化到文件,减少文件IO次数

3. 修复代码错误与冗余逻辑

  • 修正异常处理逻辑,正确判断空数据并切换备用接口
  • 移除不必要的嵌套判断,简化数据提取流程
  • 批量写入Checkpoint,避免每次处理一个条目就打开文件

优化后的代码

import datetime
import ijson
import requests

# 配置项
feed_url = "你的主接口URL"
sub_url = "你的备用接口URL"
checkpoint_file = "checkpoint.txt"
# 内存缓存已处理的job_id,避免频繁读文件
processed_job_ids = set()
# 当日时间范围,每天刷新一次
today_start = datetime.datetime.now().replace(hour=0, minute=0, second=0, microsecond=0)
today_end = today_start + datetime.timedelta(days=1)

def init_checkpoint():
    """初始化内存中的checkpoint,仅程序启动时执行一次"""
    global processed_job_ids
    try:
        with open(checkpoint_file, 'r') as f:
            processed_job_ids = set(f.read().strip().split('\n'))
    except FileNotFoundError:
        processed_job_ids = set()

def batch_save_checkpoint():
    """批量写入checkpoint,避免频繁IO"""
    with open(checkpoint_file, 'w') as f:
        f.write('\n'.join(processed_job_ids))

def is_today(creation_date_str):
    """判断日期是否为当日"""
    try:
        dt = datetime.datetime.strptime(creation_date_str, '%d.%m.%Y %H:%M:%S')
        return today_start <= dt < today_end
    except ValueError:
        return False

def get_feed_items():
    global processed_job_ids
    items = []
    response = None

    # 尝试主接口,带缓存头
    try:
        headers = {}
        # 如果之前下载过,带上Last-Modified头
        if hasattr(get_feed_items, 'last_modified'):
            headers['If-Modified-Since'] = get_feed_items.last_modified

        response = requests.get(feed_url, headers=headers, stream=True)
        # 304表示数据未更新,直接返回空
        if response.status_code == 304:
            return items

        # 更新Last-Modified缓存
        if 'Last-Modified' in response.headers:
            get_feed_items.last_modified = response.headers['Last-Modified']

        # 流式解析JSON,遍历每个job条目
        parser = ijson.items(response.raw, 'item')
        for item in parser:
            job = item.get('job', {})
            job_id = job.get('id')
            if not job_id or job_id in processed_job_ids:
                continue

            # 筛选条件:status==6 且 当日创建
            if job.get('status') == 6 and is_today(job.get('creationDate', '')):
                try:
                    long_name = job['name']
                    name = long_name.split('_')[1]
                    # 提取history数据
                    for history_item in item.get('tasks', [{}])[0].get('history', []):
                        item_data = {
                            'name': name,
                            'job_id': job_id,
                            'renderer': job.get('renderer'),
                            'status': job['status'],
                            'comment': history_item.get('comment', ''),
                        }
                        items.append(item_data)
                    # 标记为已处理
                    processed_job_ids.add(job_id)
                except KeyError as e:
                    print(f"数据格式错误,跳过条目 {job_id}: {e}")
                    continue
    except requests.exceptions.RequestException as e:
        print(f"主接口请求失败,切换备用接口: {e}")
        # 备用接口逻辑(复用筛选逻辑)
        try:
            response = requests.get(sub_url, stream=True)
            parser = ijson.items(response.raw, 'item')
            for item in parser:
                job = item.get('job', {})
                job_id = job.get('id')
                if not job_id or job_id in processed_job_ids:
                    continue
                if job.get('status') == 6 and is_today(job.get('creationDate', '')):
                    try:
                        long_name = job['name']
                        name = long_name.split('_')[1]
                        for history_item in item.get('tasks', [{}])[0].get('history', []):
                            item_data = {
                                'name': name,
                                'job_id': job_id,
                                'renderer': job.get('renderer'),
                                'status': job['status'],
                                'comment': history_item.get('comment', ''),
                            }
                            items.append(item_data)
                        processed_job_ids.add(job_id)
                    except KeyError as e:
                        print(f"备用接口数据格式错误,跳过条目 {job_id}: {e}")
                        continue
        except requests.exceptions.RequestException as e:
            print(f"备用接口请求失败: {e}")
    finally:
        if response:
            response.close()
    
    # 批量写入checkpoint(可根据执行频率调整写入时机)
    batch_save_checkpoint()
    return items

# 初始化checkpoint
init_checkpoint()
# 初始化Last-Modified属性
get_feed_items.last_modified = None

额外建议

  • 如果服务器支持自定义筛选参数(比如通过URL参数指定status=6和creationDate=today),优先使用接口筛选,这是效率最高的方案
  • 可以定时刷新today_start和today_end,避免程序运行跨天时日期判断错误
  • 对于极高频率的执行(每5秒一次),可以考虑把Checkpoint放在内存数据库中,进一步提升读写速度

内容的提问来源于stack exchange,提问作者Saba Guzitashvili

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.15 18:43:22