You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Avro Python从CSV导入报错:avro.io.AvroTypeException问题排查

解决Avro解析CSV时的avro.io.AvroTypeException问题

嘿,我来帮你排查下问题所在!你遇到的avro.io.AvroTypeException主要有两个核心原因,咱们一步步拆解解决:

1. CSV文件格式错误

你提供的CSV内容把表头和数据挤在了同一行,这会导致csv.DictReader无法正确识别字段与对应值的映射关系。正确的CSV格式应该是表头单独占一行,数据在下一行:

TransactionId,Id
2018040101000222749,1

2. 数据类型不匹配

csv.DictReader读取CSV时,所有字段的值默认都会被解析成字符串类型,但你的Avro Schema里Id字段明确定义为int类型。当你直接把row字典传给AvroProducer.produce()时,Avro的类型校验机制会发现Id的类型不符合Schema要求,从而抛出异常。

修正后的Python代码

你需要在发送数据前,手动把Id字段从字符串转换为整数:

from confluent_kafka import avro
from confluent_kafka.avro import AvroProducer
import csv

value_schema = avro.load('/home/daniela/avro/example.avsc')
AvroProducerConf = {
    'bootstrap.servers': 'localhost:9092',
    'schema.registry.url': 'http://localhost:8081',
}

avroProducer = AvroProducer(AvroProducerConf, default_value_schema=value_schema)

with open('/home/usertest/avro/data/paymenttransactions.csv') as file:
    reader = csv.DictReader(file, delimiter=",")
    for row in reader:
        # 关键步骤:将Id字段转换为整数类型
        row['Id'] = int(row['Id'])
        avroProducer.produce(topic='test', value=row)
        print(row)
    avroProducer.flush()

额外注意事项

如果你的CSV数据里存在多余空格(比如字段值前后有空格),建议在转换前先做去除处理,避免转换失败:

row['Id'] = int(row['Id'].strip())
row['TransactionId'] = row['TransactionId'].strip()

这样调整后,应该就能成功把CSV数据转换成符合Avro Schema的格式并发送到Kafka了!

内容的提问来源于stack exchange,提问作者ilpupina

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 07:25:27