You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

YOLOv5/YOLOv7目标检测JSON读取报错及Colab内存问题求助

解决YOLOv5/YOLOv7读取JSON标注报错“Expecting value: line 1 column 1 (char 0)”的方案

一、排查并修复JSON文件问题

这个报错核心原因多为JSON文件本身不合法,按以下步骤处理:

  • 批量检测JSON合法性:
    编写小脚本快速排查少量样本(比如前100个),定位损坏文件:
    import json
    import os
    
    def check_json_validity(json_dir):
        for filename in os.listdir(json_dir)[:100]:
            if filename.endswith('.json'):
                file_path = os.path.join(json_dir, filename)
                try:
                    with open(file_path, 'r') as f:
                        json.load(f)
                except Exception as e:
                    print(f"无效JSON文件: {file_path}, 错误信息: {str(e)}")
    
    check_json_validity('/你的JSON文件夹路径')
    
  • 定位空文件或损坏文件:
    筛选空文件或开头非合法JSON字符的损坏文件:
    import os
    
    def find_invalid_files(json_dir):
        for filename in os.listdir(json_dir):
            if filename.endswith('.json'):
                file_path = os.path.join(json_dir, filename)
                if os.path.getsize(file_path) == 0:
                    print(f"空JSON文件: {file_path}")
                else:
                    with open(file_path, 'r') as f:
                        first_char = f.read(1)
                        if first_char not in ['{', '[']:
                            print(f"疑似损坏的JSON文件: {file_path}, 首字符: {repr(first_char)}")
    
    find_invalid_files('/你的JSON文件夹路径')
    
  • 修复或清理问题文件:
    • 空文件直接删除;
    • 若为编码问题,尝试用utf-8-sig编码读取(修改open参数为encoding='utf-8-sig');
    • 损坏严重且无备份的文件,对应的图片可从数据集中剔除。

二、解决Colab内存不足问题

51万条数据一次性加载必然导致内存溢出,按以下方式优化:

  • 切换Colab高资源运行时:
    点击Colab顶部「运行时」→「更改运行时类型」,选择GPU作为硬件加速器,并勾选「高RAM」选项,获取更大内存空间。
  • 分批读取处理数据:
    用生成器逐批加载JSON,避免一次性占满内存:
    import json
    import os
    import gc
    
    def json_batch_generator(json_dir, batch_size=1000):
        batch = []
        for filename in os.listdir(json_dir):
            if filename.endswith('.json'):
                file_path = os.path.join(json_dir, filename)
                try:
                    with open(file_path, 'r') as f:
                        data = json.load(f)
                        batch.append(data)
                        if len(batch) == batch_size:
                            yield batch
                            batch = []
                except Exception as e:
                    print(f"跳过无效文件: {file_path}, 错误信息: {str(e)}")
        if batch:
            yield batch
    
    # 逐批处理标注数据(示例:转成YOLO格式)
    for batch_data in json_batch_generator('/你的JSON文件夹路径'):
        # 这里写你的批量处理逻辑,比如convert_to_yolo(batch_data)
        convert_to_yolo(batch_data)
        # 释放当前批次内存
        del batch_data
        gc.collect()
    
  • 清理冗余内存占用:
    每批数据处理完成后,用del删除变量并调用gc.collect()强制回收内存,避免内存泄漏。

内容的提问来源于stack exchange,提问作者ewise

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 15:05:30