Databricks 12.2集成Confluent Schema Registry时出现Schema Not Found错误
问题排查与解决方案
核心问题:subject参数类型错误
你在to_avro中使用F.lit("demo-topic-value")作为subject参数,这是错误用法。subject参数要求传入字符串常量,而非Column类型。当传入Column表达式时,Spark会在任务执行阶段尝试解析该表达式,导致Schema Registry客户端无法正确识别目标subject名称,最终返回40403错误。
修正代码
将subject参数改为直接传入字符串常量:
df_tpch_orders.withColumn("value",to_avro( data = df_tpch_orders.o_comment, options = schema_registry_options, schemaRegistryAddress = schema_registry_address, subject = "demo-topic-value", # 替换为字符串常量 ))
其他可能的排查点
如果修正后仍有问题,检查以下内容:
- Schema Registry配置一致性:确认
schema_registry_options中包含必要的认证参数(如basic.auth.user.info),且与你手动验证schema时的配置完全一致。部分场景下,认证信息缺失或错误会伪装成“Schema未找到”的404错误。 - 依赖兼容性:Databricks 12.2基于Spark 3.3.x,确保环境中
spark-avro与Confluent Schema Registry客户端版本兼容,避免因依赖冲突导致客户端无法正常请求Schema Registry。 - Subject名称大小写:确认
demo-topic-value的大小写与Schema Registry中实际存在的subject完全一致,Schema Registry的subject名称是大小写敏感的。
内容的提问来源于stack exchange,提问作者Raviteja Sutrave
相关产品推荐
相关产品推荐

