运行infer_detections脚本时出现UnicodeDecodeError错误求助
问题排查:运行infer_detections脚本时的UnicodeDecodeError
问题重现
执行以下检测推理命令时,抛出了UnicodeDecodeError编码错误:
python -m infer_detections --input_tfrecord_paths=../data/coco_testdev.record --output_tfrecord_path=../data/inference --inference_graph=../model/fine_tuned_model/frozen_inference_graph.pb --discard_image_pixels
错误核心信息:
UnicodeDecodeError: 'utf-8' codec can't decode byte 0xff in position 394: invalid start byte
Traceback (most recent call last): File "C:\ProgramData...
可能的原因及解决方案
1. TFRecord中包含非UTF-8编码的字符串字段
TFRecord里的元数据(比如文件名、标注文本)大概率用了GBK、Latin-1这类非UTF-8编码,而脚本默认以UTF-8解码,就会触发这个错误。
解决步骤:
- 先写个小脚本排查TFRecord字段的编码情况:
import tensorflow as tf def check_tfrecord_encoding(tfrecord_path): raw_dataset = tf.data.TFRecordDataset(tfrecord_path) # 检查前5条记录足够定位问题 for idx, raw_record in enumerate(raw_dataset.take(5)): print(f"\n=== 第{idx+1}条记录 ===") example = tf.train.Example() example.ParseFromString(raw_record.numpy()) for key, feature in example.features.feature.items(): if feature.HasField('bytes_list') and feature.bytes_list.value: byte_data = feature.bytes_list.value[0] try: decoded = byte_data.decode('utf-8') print(f"字段 {key}: {decoded[:50]}...") except UnicodeDecodeError: # 尝试常见的替代编码 for enc in ['gbk', 'latin-1']: try: decoded = byte_data.decode(enc) print(f"字段 {key} ({enc}解码): {decoded[:50]}...") break except: continue else: print(f"字段 {key}: 无法用utf-8/gbk/latin-1解码,字节开头: {byte_data[:10]}") check_tfrecord_encoding("../data/coco_testdev.record") - 根据排查结果修改
infer_detections脚本的解码逻辑:
如果是编码不兼容,要么把解码换成对应编码,要么添加错误忽略:# 原代码示例 # filename = feature['filename'].bytes_list.value[0].decode('utf-8') # 方案1:换成正确编码(比如gbk) filename = feature['filename'].bytes_list.value[0].decode('gbk') # 方案2:忽略解码错误(不推荐,但应急可用) filename = feature['filename'].bytes_list.value[0].decode('utf-8', errors='ignore')
2. 脚本错误地将二进制数据当作字符串解码
虽然你加了--discard_image_pixels参数,但脚本可能仍在处理某些二进制字段时,错误地调用了UTF-8解码方法。
解决步骤:
- 从完整堆栈跟踪里找到报错的具体代码行,确认是哪个字段出了问题。如果是纯二进制数据(比如非文本的字节字段),直接去掉
decode()调用,保留原始字节数据即可。
3. TFRecord文件损坏
如果TFRecord在生成、下载或传输过程中损坏,也会出现解码异常。
解决步骤:
- 重新生成
coco_testdev.record文件,确保生成过程没有中断;如果是下载的文件,重新下载并校验文件完整性。
内容的提问来源于stack exchange,提问作者Oblomov
相关产品推荐
相关产品推荐

