如何在图像文本匹配模型中计算corpus_bleu分数
问题描述
已训练好基于BERT的文本编码器模型,需要在预测数据上计算corpus_bleu分数,但不清楚该在代码中何处实现。这是一个图像字幕匹配模型,训练了图像编码器和文本编码器,现有代码如下:
加载模型
vision_encoder = keras.models.load_model("vision_encoder") text_encoder = keras.models.load_model("text_encoder")
数据读取函数
def read_image(image_path): image_array = tf.image.decode_jpeg(tf.io.read_file(image_path), channels=3) return tf.image.resize(image_array, (299, 299)) def read_text(caption): return tf.convert_to_tensor(caption)
加载数据集
j = df['findings'].astype(str) j.shape # 输出: (3851,)
生成文本嵌入
text_embeddings = text_encoder.predict( tf.data.Dataset.from_tensor_slices(j) .map(read_text).batch(batch_size), verbose=1, ) print(f"Text embeddings shape: {text_embeddings.shape}.") # 输出: 1926/1926 [==============================] - 25s 13ms/step # Text embeddings shape: (3851, 128, 256).
匹配函数
def find_matches(t_embeddings, queries, k=9, normalize=True): results_list = [] for image_path in queries: image_array = tf.image.decode_jpeg(tf.io.read_file(image_path), channels=3) imgr = tf.expand_dims(image_array, axis=0) i_embedding = vision_encoder(tf.image.resize(imgr, (299, 299))) # 归一化 if normalize: image_embeddings = tf.math.l2_normalize(t_embeddings, axis=1) query_embedding = tf.math.l2_normalize(i_embedding, axis=1) # 计算相似度 dot_similarity = tf.matmul(query_embedding, image_embeddings, transpose_b=True) # 获取top k索引 results = tf.math.top_k(dot_similarity, k).indices.numpy()[0] # 获取对应字幕 matched_captions = [df['findings'][idx] for idx in results] results_list.append(matched_captions) return results_list
现有匹配调用
img = "/content/image.png" matches = find_matches(t_embeddings, [img], normalize=True)[0] for i in range(9): print(matches[i])
实现
corpus_bleu的位置与代码 corpus_bleu需要两组核心数据:每个测试样本的真实字幕集合和模型预测的字幕集合,因此要在批量处理测试图像、获取匹配字幕之后执行计算,具体步骤如下:
1. 准备依赖
首先导入所需工具库:
import nltk from nltk.translate.bleu_score import corpus_bleu nltk.download('punkt')
如果处理中文字幕,需替换分词工具为jieba等中文分词库
2. 整理测试数据集
假设你有测试集存储在test_df中,包含image_path(图像路径)和findings(对应真实字幕)字段:
# 提取测试图像路径与真实字幕 test_image_paths = test_df['image_path'].tolist() # corpus_bleu要求真实字幕为「列表的列表」结构,每个子列表对应一个样本的所有可能真实字幕 true_captions = [[caption.split()] for caption in test_df['findings'].tolist()]
3. 批量获取预测字幕
遍历所有测试图像,调用find_matches获取每个图像的Top1匹配字幕(作为模型预测结果):
predicted_captions = [] # 批量获取所有测试图像的匹配字幕 matches_list = find_matches(text_embeddings, test_image_paths, k=1, normalize=True) # 提取每个图像的Top1字幕并转为单词列表 for matches in matches_list: pred_caption = matches[0].split() predicted_captions.append(pred_caption)
4. 计算corpus_bleu分数
在收集完所有真实字幕和预测字幕后,直接调用计算:
# 默认计算BLEU-4(n-gram权重为(0.25,0.25,0.25,0.25)),可自定义权重调整计算BLEU-1/2/3 bleu_score = corpus_bleu(true_captions, predicted_captions) print(f"Corpus BLEU score: {bleu_score:.4f}")
关键注意事项
- 原
find_matches函数存在变量未定义问题(未从queries中读取image_path),已在上述代码中修正 - 若每个图像对应多个真实字幕,只需在
true_captions的子列表中添加更多分词后的字幕即可
内容的提问来源于stack exchange,提问作者user19686684
相关产品推荐
相关产品推荐

