You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python打开json.gz.part分割压缩文件?

解决拆分的gzip.part文件解压及JSON解析问题

问题说明

目录里有大量json.gz文件,其中部分是json.gz.part拆分文件(推测是大文件保存时被拆分)。尝试常规gzip读取方法时触发解压错误:

error: Error -3 while decompressing data: invalid code lengths set

尝试jiffyclub的解决方案时又遇到属性缺失错误:

AttributeError: type object 'GzipFile' has no attribute '_read_eof'

补充:逐行读取能获取大部分内容,但最终还是会触发相同的解压错误,且无法直接解析为完整JSON。

可行解决方案

1. 优先合并拆分的gzip文件

.part本质是未完成的拆分压缩包,先把同组的.part文件按顺序合并,再处理才是最稳妥的:

  • 先确认同组文件的命名规则(比如data.json.gz.part001、data.json.gz.part002),按数字排序
  • 用命令行合并:
    # Linux/macOS系统
    cat data.json.gz.part* > merged_data.json.gz
    
    # Windows系统
    copy /b data.json.gz.part001 + data.json.gz.part002 merged_data.json.gz
    
  • 合并完成后用常规方法读取:
    import gzip
    import json
    
    with gzip.open('merged_data.json.gz', 'r') as fin:
        json_data = json.load(fin)
    

2. 无法合并时,读取gzip.part的可用内容

如果部分.part文件丢失,只能尝试读取未损坏的部分并忽略错误:

import gzip
import json
from io import BytesIO

def read_partial_gzip(file_path):
    buffer = BytesIO()
    with open(file_path, 'rb') as f:
        while True:
            chunk = f.read(1024)
            if not chunk:
                break
            try:
                # 尝试解压当前块
                decompressor = gzip.GzipFile(fileobj=BytesIO(chunk))
                buffer.write(decompressor.read())
            except Exception as e:
                # 遇到损坏块直接跳过,继续读后续内容
                print(f"跳过损坏块: {e}")
                continue
    # 重置缓冲区指针,尝试解析JSON
    buffer.seek(0)
    try:
        return json.load(buffer)
    except json.JSONDecodeError:
        # 如果JSON不完整,截取到最后一个有效闭合括号
        buffer.seek(0)
        content = buffer.read().decode('utf-8')
        last_brace = content.rfind('}')
        if last_brace != -1:
            valid_content = content[:last_brace+1]
            return json.loads(valid_content)
        else:
            raise ValueError("未找到有效JSON内容")

# 使用示例
data = read_partial_gzip('your_file.json.gz.part')

3. 关于_read_eof属性缺失的说明

jiffyclub的方案是针对旧版Python的gzip模块设计的,新版Python已经移除了_read_eof这个私有属性,不用再纠结这个方案,直接用上面的合并或部分读取方法就行。

额外提示

如果你的JSON是多行格式(每行一个独立JSON对象),逐行读取时可以提前收集有效行:

import gzip
import json

valid_data = []
with gzip.open('file.json.gz.part', 'r') as fin:
    for line in fin:
        try:
            obj = json.loads(line.decode('utf-8'))
            valid_data.append(obj)
        except Exception as e:
            print(f"跳过无效行: {e}")
            break

内容的提问来源于stack exchange,提问作者Marlon Teixeira

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 02:31:08