You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何读取含JSON内容的gzip文件并写入文本文件,解决bytes不可JSON序列化报错

问题根因
  • 错误调用json.dumps():你的核心需求是解压gzip文件并写入文本,不需要执行JSON序列化操作,这行是多余代码
  • 类型不匹配报错:gzip.open以rb模式读取返回的是bytes类型字节流,json.dumps()仅支持处理字典、列表、字符串等可JSON序列化的Python对象,无法直接序列化bytes数据
  • 逻辑缺失:现有代码未将解压后的内容写入输出文件,json.dumps()的返回值没有被接收使用,即使无报错也无法得到预期结果
修复方案

场景1:仅需解压gzip内容写入文本,无需处理JSON

直接用文本模式读取gzip文件,跳过所有JSON相关操作即可,代码如下:

import gzip

# rt为gzip文本读取模式,自动完成解压+字节转字符串操作
with gzip.open(".../2020-04/statuses.log.2020-04-01-00.gz", 'rt', encoding='utf-8') as f_in:
    with open('.../notebooks/decompressed.txt', 'w', encoding='utf-8') as f_out:
        f_out.write(f_in.read())

大文件可以逐行写入降低内存占用:

import gzip

with gzip.open(".../2020-04/statuses.log.2020-04-01-00.gz", 'rt', encoding='utf-8') as f_in:
    with open('.../notebooks/decompressed.txt', 'w', encoding='utf-8') as f_out:
        for line in f_in:
            f_out.write(line)

场景2:需要解析/处理文件内的JSON内容再写入

如果文件是每行一个独立JSON对象的NDJSON格式(常见日志结构),可以逐行解析处理后再写入:

import gzip
import json

with gzip.open(".../2020-04/statuses.log.2020-04-01-00.gz", 'rt', encoding='utf-8') as f_in:
    with open('.../notebooks/decompressed.txt', 'w', encoding='utf-8') as f_out:
        for line in f_in:
            line = line.strip()
            if not line:
                continue
            # 解析单行JSON
            json_data = json.loads(line)
            # 此处可添加对json_data的处理逻辑
            # 转成字符串写入文件,ensure_ascii=False避免中文乱码
            f_out.write(json.dumps(json_data, ensure_ascii=False) + '\n')

内容的提问来源于stack exchange,提问作者Trishala Suryavanshi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.01 04:36:06