You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何仅用Python 3代码实现fastavro Shell命令的功能?

纯Python实现fastavro命令的功能

导出AVRO Schema为JSON文件

对应原Shell命令:fastavro --schema -p {avro_file} > {schout_file}

使用fastavro.schema.load_schema加载AVRO文件的Schema,再通过json.dump格式化输出到文件:

import fastavro
import json

def export_avro_schema(avro_file_path, output_schema_path):
    # 加载AVRO文件的Schema
    schema = fastavro.schema.load_schema(avro_file_path)
    # 以格式化JSON写入输出文件(对应原命令的-p/--pretty参数)
    with open(output_schema_path, 'w', encoding='utf-8') as f:
        json.dump(schema, f, indent=2)

导出AVRO Records为JSON文件

对应原Shell命令:fastavro -p {avro_file} > {rcdout_file}

使用fastavro.reader读取AVRO文件中的所有记录,转成列表后格式化输出为JSON:

def export_avro_records(avro_file_path, output_records_path):
    # 读取AVRO文件中的所有记录
    with open(avro_file_path, 'rb') as f:
        reader = fastavro.reader(f)
        records = list(reader)
    # 以格式化JSON写入输出文件(对应原命令的-p/--pretty参数)
    with open(output_records_path, 'w', encoding='utf-8') as f:
        json.dump(records, f, indent=2)

调用示例

# 替换为你的文件路径
avro_file = "input.avro"
schout_file = "schema.json"
rcdout_file = "records.json"

export_avro_schema(avro_file, schout_file)
export_avro_records(avro_file, rcdout_file)

补充:大文件处理优化

如果处理超大AVRO文件,不想一次性加载所有记录到内存,可以逐行写入JSON数组,避免内存占用过高:

def export_large_avro_records(avro_file_path, output_records_path):
    with open(avro_file_path, 'rb') as f_in, open(output_records_path, 'w', encoding='utf-8') as f_out:
        reader = fastavro.reader(f_in)
        f_out.write('[\n')
        first_record = True
        for record in reader:
            if not first_record:
                f_out.write(',\n')
            json.dump(record, f_out, indent=2)
            first_record = False
        f_out.write('\n]')

内容的提问来源于stack exchange,提问作者Tarlak333

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.12 03:58:23