如何仅用Python 3代码实现fastavro Shell命令的功能?
纯Python实现fastavro命令的功能
导出AVRO Schema为JSON文件
对应原Shell命令:fastavro --schema -p {avro_file} > {schout_file}
使用fastavro.schema.load_schema加载AVRO文件的Schema,再通过json.dump格式化输出到文件:
import fastavro import json def export_avro_schema(avro_file_path, output_schema_path): # 加载AVRO文件的Schema schema = fastavro.schema.load_schema(avro_file_path) # 以格式化JSON写入输出文件(对应原命令的-p/--pretty参数) with open(output_schema_path, 'w', encoding='utf-8') as f: json.dump(schema, f, indent=2)
导出AVRO Records为JSON文件
对应原Shell命令:fastavro -p {avro_file} > {rcdout_file}
使用fastavro.reader读取AVRO文件中的所有记录,转成列表后格式化输出为JSON:
def export_avro_records(avro_file_path, output_records_path): # 读取AVRO文件中的所有记录 with open(avro_file_path, 'rb') as f: reader = fastavro.reader(f) records = list(reader) # 以格式化JSON写入输出文件(对应原命令的-p/--pretty参数) with open(output_records_path, 'w', encoding='utf-8') as f: json.dump(records, f, indent=2)
调用示例
# 替换为你的文件路径 avro_file = "input.avro" schout_file = "schema.json" rcdout_file = "records.json" export_avro_schema(avro_file, schout_file) export_avro_records(avro_file, rcdout_file)
补充:大文件处理优化
如果处理超大AVRO文件,不想一次性加载所有记录到内存,可以逐行写入JSON数组,避免内存占用过高:
def export_large_avro_records(avro_file_path, output_records_path): with open(avro_file_path, 'rb') as f_in, open(output_records_path, 'w', encoding='utf-8') as f_out: reader = fastavro.reader(f_in) f_out.write('[\n') first_record = True for record in reader: if not first_record: f_out.write(',\n') json.dump(record, f_out, indent=2) first_record = False f_out.write('\n]')
内容的提问来源于stack exchange,提问作者Tarlak333
相关产品推荐
相关产品推荐

