You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中嵌套对象的高效单行存储格式及对应模块推荐

适合嵌套对象的高效单行存储方案

现成格式推荐

  • 自定义Schema头的TSV/CSV:这是最轻量化的方案——首行用结构化字符串定义字段类型与嵌套关系(比如id(int),name(str),items.name(str)[],items.price(int)[]),后续每行用分隔符存储扁平化后的用户数据,既保留了文本可读性,又避免了JSON的字段名冗余。
  • Apache Avro:原生支持Schema驱动的存储,Schema可以内嵌在文件头部,数据行采用紧凑编码(支持二进制或可读JSON格式),空间效率远高于普通JSON,同时天然支持逐行读写与追加操作。

Python实现方案与模块

自定义TSV/CSV(灵活易上手)

用Python内置的csv模块即可实现,只需自己编写嵌套对象的扁平化与还原逻辑:

import csv
from typing import Dict, List

# 从文件首行读取或预定义Schema
SCHEMA = ["id(int)", "name(str)", "items.name(str)[]", "items.price(int)[]"]

def flatten_user(user: Dict) -> List:
    """将User嵌套对象扁平化为列表"""
    flat_data = [user["id"], user["name"]]
    # 展开items的name和price列表
    flat_data.extend(item["name"] for item in user["items"])
    flat_data.extend(item["price"] for item in user["items"])
    return flat_data

def unflatten_row(row: List, schema: List) -> Dict:
    """将扁平行数据还原为User对象"""
    user = {"id": int(row[0]), "name": row[1], "items": []}
    # 定位嵌套字段的起始索引
    name_start_idx = schema.index("items.name(str)[]")
    price_start_idx = schema.index("items.price(int)[]")
    # 提取并配对商品名称与价格
    item_names = row[name_start_idx:price_start_idx]
    item_prices = list(map(int, row[price_start_idx:]))
    for name, price in zip(item_names, item_prices):
        user["items"].append({"name": name, "price": price})
    return user

# 写入初始数据
with open("users.tsv", "w", newline="") as f:
    writer = csv.writer(f, delimiter="\t")
    writer.writerow(SCHEMA)
    writer.writerow(flatten_user({
        "id": 1,
        "name": "Alice",
        "items": [{"name": "apple", "price": 10}, {"name": "banana", "price": 20}]
    }))

# 追加新用户
with open("users.tsv", "a", newline="") as f:
    writer = csv.writer(f, delimiter="\t")
    writer.writerow(flatten_user({
        "id": 2,
        "name": "Bob",
        "items": [{"name": "orange", "price": 15}]
    }))

# 逐行读取解析
with open("users.tsv", "r") as f:
    reader = csv.reader(f, delimiter="\t")
    schema = next(reader)
    for row in reader:
        print(unflatten_row(row, schema))

Apache Avro(高效且规范)

使用fastavro模块(性能优于官方Avro库),它会自动处理Schema存储与数据编码,无需手动扁平化:

from fastavro import writer, reader, parse_schema

# 定义Avro Schema
USER_SCHEMA = parse_schema({
    "type": "record",
    "name": "User",
    "fields": [
        {"name": "id", "type": "int"},
        {"name": "name", "type": "string"},
        {"name": "items", "type": {
            "type": "array",
            "items": {
                "type": "record",
                "name": "Item",
                "fields": [
                    {"name": "name", "type": "string"},
                    {"name": "price", "type": "int"}
                ]
            }
        }}
    ]
})

# 写入初始数据
with open("users.avro", "wb") as f:
    writer(f, USER_SCHEMA, [
        {"id": 1, "name": "Alice", "items": [{"name": "apple", "price": 10}]},
        {"id": 2, "name": "Bob", "items": [{"name": "banana", "price": 20}]}
    ])

# 追加新用户
with open("users.avro", "ab") as f:
    writer(f, USER_SCHEMA, [{"id": 3, "name": "Charlie", "items": [{"name": "orange", "price": 15}]}])

# 逐行读取
with open("users.avro", "rb") as f:
    for user in reader(f):
        print(user)

如果需要文本格式的可读数据,可在写入时指定format="json",空间效率仍优于普通JSON Lines。

其他可选方案

  • Parquet:通过pyarrow模块实现,空间压缩比极高,适合大数据量场景,但操作复杂度略高,小项目无需考虑。
  • 完全自定义格式:用JSON在首行描述Schema,后续行用分隔符存储值,自己编写解析逻辑,灵活性最高但需要维护更多代码。

内容的提问来源于stack exchange,提问作者Vince M

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 21:20:16