You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用apache_avro crate时,如何在顶层Schema中正确引用子Schema?

问题

我尝试用Schema::parse_list方法编写嵌套记录(让顶层Schema引用子Schema),但向Writer追加数据时触发错误。直接在顶层Schema中嵌套定义字段的方式能正常运行,请问怎么在顶层Schema里正确引用子Schema?

代码示例

use std::fs::{File, OpenOptions};
use std::io::Write;
use apache_avro::{Schema, Writer};
use apache_avro::types::Record;

fn write_data (top_schema: &Schema, sub_schema: &Schema) {
    let mut sub_record = Record::new(&sub_schema).unwrap();
    sub_record.put("streetaddress", "123 Main St.");
    sub_record.put("city", "Anytown");

    let mut top_record = Record::new(&top_schema).unwrap();
    top_record.put("firstname", "John");
    top_record.put("lastname", "Doe");
    top_record.put("address", sub_record);


    let mut file = OpenOptions::new()
        .write(true)
        .create(true)
        .open("customer.avro")
        .unwrap();

    let mut writer = Writer::new(&top_schema, &mut file);
    writer.append(top_record).unwrap();
    writer.flush().unwrap();
}

fn create_data_without_schema_references(){
    let top_schema = Schema::parse_str(r#"
    {
        "name": "Customer",
        "type": "record",
        "fields": [
            {"name": "firstname", "type": "string"},
            {"name": "lastname", "type": "string"},
            {"name": "address", "type":
                {
                    "name": "AddressRecord",
                    "type": "record",
                    "fields": [
                    {"name": "streetaddress", "type": "string"},
                    {"name": "city", "type": "string"}
                    ]}
            }
        ]
    }
    "#).unwrap();

    let sub_schema = Schema::parse_str(r#"
    {
        "name": "AddressRecord",
        "type": "record",
        "fields": [
            {"name": "streetaddress", "type": "string"},
            {"name": "city", "type": "string"}
        ]
    }
    "#).unwrap();

    write_data(&top_schema, &sub_schema);
}

fn create_data_with_schema_reference(){

    let top_schema = r#"
    {
        "name": "Customer",
        "type": "record",
        "fields": [
            {"name": "firstname", "type": "string"},
            {"name": "lastname", "type": "string"},
            {"name": "address", "type":"AddressRecord"}
        ]
    }
    "#;

    let sub_schema = r#"
    {
        "name": "AddressRecord",
        "type": "record",
        "fields": [
            {"name": "streetaddress", "type": "string"},
            {"name": "city", "type": "string"}
        ]
    }
    "#;

    let schemas = Schema::parse_list(&[sub_schema, top_schema]).unwrap();

    write_data(&schemas[1], &schemas[0]);
}

fn main() {
    create_data_without_schema_references(); // 正常运行
    create_data_with_schema_reference(); // 运行报错
}

错误信息

thread 'main' panicked at 'called `Result::unwrap()` on an `Err` value: SchemaResolutionError(Name { name: "AddressRecord", namespace: None })', src/main.rs:30:31
解决方案

问题原因

Schema::parse_list解析得到的两个Schema实例是独立关联的:顶层Schema在解析时已经通过引用绑定了内部的AddressRecord定义,但传入write_data的子Schema是另一个独立实例。用这个独立子Schema创建Record后,顶层Schema无法识别该Record的类型,导致Schema解析错误。

修复方式

不需要单独传入子Schema,而是从顶层Schema中直接获取嵌套字段对应的Schema来创建子Record。因为Schema::parse_list已经按顺序解析了子Schema和顶层Schema,两者的引用关系已经建立完成。

修改后的代码如下:

调整write_data函数

fn write_data(top_schema: &Schema) {
    // 从顶层Schema中获取address字段对应的Schema
    let address_schema = top_schema.get_field("address")
        .unwrap()
        .schema();
    
    let mut sub_record = Record::new(address_schema).unwrap();
    sub_record.put("streetaddress", "123 Main St.");
    sub_record.put("city", "Anytown");

    let mut top_record = Record::new(top_schema).unwrap();
    top_record.put("firstname", "John");
    top_record.put("lastname", "Doe");
    top_record.put("address", sub_record);

    let mut file = OpenOptions::new()
        .write(true)
        .create(true)
        .open("customer_ref.avro")
        .unwrap();

    let mut writer = Writer::new(top_schema, &mut file);
    writer.append(top_record).unwrap();
    writer.flush().unwrap();
}

调整create_data_with_schema_reference函数

fn create_data_with_schema_reference(){
    let top_schema_str = r#"
    {
        "name": "Customer",
        "type": "record",
        "fields": [
            {"name": "firstname", "type": "string"},
            {"name": "lastname", "type": "string"},
            {"name": "address", "type":"AddressRecord"}
        ]
    }
    "#;

    let sub_schema_str = r#"
    {
        "name": "AddressRecord",
        "type": "record",
        "fields": [
            {"name": "streetaddress", "type": "string"},
            {"name": "city", "type": "string"}
        ]
    }
    "#;

    let schemas = Schema::parse_list(&[sub_schema_str, top_schema_str]).unwrap();
    write_data(&schemas[1]);
}

修复原理

Schema::parse_list会按传入顺序解析Schema,先解析AddressRecord,这样顶层Schema在解析时能直接找到该类型的定义并建立内部引用。此时从顶层Schema的address字段获取的Schema,就是解析后的AddressRecord实例,用它创建的Record和顶层Schema的类型要求完全匹配,Writer就能正常处理。

内容的提问来源于stack exchange,提问作者kh46r

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 04:10:36