带引用链的JSON Schema处理及Nifi处理器开发方案咨询
Hey there, let's work through your JSON Schema challenges step by step—you're already on the right track with the dereferencing idea, so let's expand that into actionable solutions for each of your tasks.
Problem Recap
To make sure we're aligned, here's your current setup and roadblocks:
- Core Tasks:
- Validate JSON data against JSON Schema
- Build a converter from JSON Schema to AVRO Schema
- Build a converter from JSON Schema to Hive Table DDL
- Current Issues:
- The
java-json-tools/json-schema-validatorthrows errors due to chained/relative$refreferences in your schema - No ready-to-use open-source libraries exist for the AVRO and Hive conversion tasks
- You've already built a NiFi Processor for task 1, and want to use an inline parser to dereference schemas into a single, complete document before proceeding.
- The
Why Dereferencing First Is Non-Negotiable
Dereferencing is the perfect starting point here. Most validation and conversion tools struggle with external or nested $ref references because they don't handle cross-file resolution out of the box. Flattening your schema into a self-contained document eliminates reference-related errors and creates a consistent input for all three tasks.
Task 1: Fix JSON Schema Validation with Dereferencing
The java-json-tools library (specifically com.github.fge:json-schema-validator) has built-in tools to resolve references—you just need to configure a schema store to load relative files. Here's how to implement it:
// Load your base schema file final JsonNode schemaNode = JsonLoader.fromFile(new File("/path/to/carrier_schema.json")); // Create a schema store pointing to the base directory of your schemas // This lets the resolver find files like ../identification_service.json final SchemaStore store = new FileSystemSchemaStore(new File("/base/path/to/all/schemas")); // Dereference the schema to flatten all $ref references final JsonSchemaDereferencer dereferencer = new DefaultDereferencer(store); final JsonNode dereferencedSchema = dereferencer.dereference(schemaNode); // Now validate against the flattened schema final JsonSchemaFactory factory = JsonSchemaFactory.byDefault(); final JsonSchema schema = factory.getJsonSchema(dereferencedSchema); final ProcessingReport report = schema.validate(jsonDataNode);
If you need to handle remote HTTP schemas later, you can extend SchemaStore to add HTTP fetching logic, but filesystem resolution works for your relative-path use case.
Task 2: JSON Schema to AVRO Schema Converter
Since no off-the-shelf library handles complex schema conversions perfectly, building a custom converter on top of your dereferenced schema is the way to go. Here's a structured approach:
- Map JSON Schema types to AVRO types:
string→stringinteger→int/long(use format hints likeint32/int64to decide)number→float/doubleboolean→booleanobject→record(mappropertiesto AVRO fields)array→array(mapitemsto the corresponding AVRO type)- Nullable fields → AVRO union types (e.g.,
["null", "string"])
- Handle schema metadata:
- Map
title/descriptionto AVRO'sdocfield - Mark
requiredfields as non-nullable (excludenullfrom the union)
- Map
- Use the Avro library to build schemas programmatically:
// Example: Convert a JSON Schema object to an AVRO Record private Schema mapJsonSchemaToAvro(JsonNode schemaNode) { if (schemaNode.get("type").asText().equals("object")) { RecordBuilder recordBuilder = SchemaBuilder.record(schemaNode.get("title").asText()) .namespace("your.avro.namespace") .doc(schemaNode.get("description").asText()); JsonNode properties = schemaNode.get("properties"); JsonNode required = schemaNode.get("required"); for (Map.Entry<String, JsonNode> prop : properties.fields()) { String fieldName = prop.getKey(); JsonNode propSchema = prop.getValue(); Schema fieldType = mapJsonSchemaToAvro(propSchema); // Handle required fields if (required != null && required.contains(fieldName)) { recordBuilder = recordBuilder.field(fieldName).type(fieldType).noDefault(); } else { recordBuilder = recordBuilder.field(fieldName) .type(SchemaBuilder.unionOf().nullType().and().type(fieldType).end()) .defaultValue(null); } } return recordBuilder.build(); } // Add logic for other types (string, integer, array, etc.) }
Task 3: JSON Schema to Hive Table DDL Converter
Again, a custom converter built on the dereferenced schema is your best bet. Follow this workflow:
- Map JSON Schema types to Hive types:
string→STRINGinteger→INT;number→DOUBLEboolean→BOOLEANobject→STRUCT(with nested fields)array→ARRAY<type>
- Handle schema metadata:
- Add
title/descriptionas Hive comments usingCOMMENT '...' - Mark
requiredfields as non-nullable (omitNULLtype or useNOT NULLfor Hive strict mode)
- Add
- Generate DDL programmatically:
-- Example DDL for your carrier schema (after dereferencing) CREATE EXTERNAL TABLE IF NOT EXISTS carrier_identification ( type STRING COMMENT 'Event type', event_data STRUCT< result: STRUCT< mno: STRING COMMENT 'Mobile network operator', mvno: STRING COMMENT 'Mobile virtual network operator', mcc: STRING COMMENT 'Mobile Country Code', mnc: STRING COMMENT 'Mobile Network Code', country: STRING COMMENT 'ISO 3166-1 alpha 2 country code' > COMMENT 'Carrier identification result' > COMMENT 'A successfully identified carrier of a user' ) LOCATION '/path/to/hive/table/storage' ROW FORMAT SERDE 'org.openx.data.jsonserde.JsonSerDe';
NiFi Processor Integration
Since you've already built the validation processor, extend it with these steps for the conversion tasks:
- Create a shared utility class for schema dereferencing—reuse this across all three processors to avoid duplicate code.
- For conversion processors:
- Add a NiFi property to let users specify the base directory for schema files (to resolve relative references)
- Fetch the input schema from a flow file, local file, or NiFi Registry
- Dereference the schema using the shared utility
- Run the conversion logic (AVRO or Hive DDL)
- Output the converted schema as a new flow file
- Add error handling: Route flow files to a failure relationship if dereferencing or conversion fails, with attributes capturing the error details.
Example Dereferenced Schema Preview
Here's a simplified version of your sample schema after dereferencing (assuming ../identification_service.json's fields are inline):
{
"id": "/schemas/bi/events/identification/carrier",
"$schema": "http://json-schema.org/draft-04/schema#",
"title": "Users Carrier Identified",
"description": "A successfully identified carrier of a user",
"type": "object",
"properties": {
"type": {
"constant": "identification.carrier",
"example": "identification.carrier"
},
"event_data": {
"type": "object",
"properties": {
"result": {
"type": "object",
"properties": {
"mno": { "type": "string", "title": "Mobile network operator", "example": "Telekom" },
"mvno": { "type": "string", "title": "Mobile virtual network operator", "example": "Mobilcom-Debitel" },
"mcc": { "type": "string", "title": "Mobile Country Code", "example": "262" },
"mnc": { "type": "string", "title": "Mobile Network Code", "example": "01" },
"country": { "type": "string", "title": "ISO 3166-1 alpha 2 country code", "example": "DE" }
},
"required": ["mno", "country"]
}
},
"required": ["result"]
}
},
"required": ["type", "event_data"]
}
内容的提问来源于stack exchange,提问作者Sam

