You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在图像文本匹配模型中计算corpus_bleu分数

问题描述

已训练好基于BERT的文本编码器模型,需要在预测数据上计算corpus_bleu分数,但不清楚该在代码中何处实现。这是一个图像字幕匹配模型,训练了图像编码器和文本编码器,现有代码如下:

加载模型

vision_encoder = keras.models.load_model("vision_encoder")
text_encoder = keras.models.load_model("text_encoder")

数据读取函数

def read_image(image_path):
    image_array = tf.image.decode_jpeg(tf.io.read_file(image_path), channels=3)
    return tf.image.resize(image_array, (299, 299))

def read_text(caption):
    return tf.convert_to_tensor(caption)

加载数据集

j = df['findings'].astype(str)
j.shape
# 输出: (3851,)

生成文本嵌入

text_embeddings = text_encoder.predict(
    tf.data.Dataset.from_tensor_slices(j)
    .map(read_text).batch(batch_size),
    verbose=1,
)
print(f"Text embeddings shape: {text_embeddings.shape}.")
# 输出: 1926/1926 [==============================] - 25s 13ms/step
# Text embeddings shape: (3851, 128, 256).

匹配函数

def find_matches(t_embeddings, queries, k=9, normalize=True):
    results_list = []
    for image_path in queries:
        image_array = tf.image.decode_jpeg(tf.io.read_file(image_path), channels=3)
        imgr = tf.expand_dims(image_array, axis=0)
        i_embedding = vision_encoder(tf.image.resize(imgr, (299, 299)))
        # 归一化
        if normalize:
            image_embeddings = tf.math.l2_normalize(t_embeddings, axis=1)
            query_embedding = tf.math.l2_normalize(i_embedding, axis=1)
        # 计算相似度
        dot_similarity = tf.matmul(query_embedding, image_embeddings, transpose_b=True)
        # 获取top k索引
        results = tf.math.top_k(dot_similarity, k).indices.numpy()[0]
        # 获取对应字幕
        matched_captions = [df['findings'][idx] for idx in results]
        results_list.append(matched_captions)
    return results_list

现有匹配调用

img = "/content/image.png"
matches = find_matches(t_embeddings, 
                       [img], 
                       normalize=True)[0]

for i in range(9):
    print(matches[i])

实现corpus_bleu的位置与代码

corpus_bleu需要两组核心数据:每个测试样本的真实字幕集合和模型预测的字幕集合,因此要在批量处理测试图像、获取匹配字幕之后执行计算,具体步骤如下:

1. 准备依赖

首先导入所需工具库:

import nltk
from nltk.translate.bleu_score import corpus_bleu
nltk.download('punkt')

如果处理中文字幕,需替换分词工具为jieba等中文分词库

2. 整理测试数据集

假设你有测试集存储在test_df中,包含image_path(图像路径)和findings(对应真实字幕)字段:

# 提取测试图像路径与真实字幕
test_image_paths = test_df['image_path'].tolist()
# corpus_bleu要求真实字幕为「列表的列表」结构,每个子列表对应一个样本的所有可能真实字幕
true_captions = [[caption.split()] for caption in test_df['findings'].tolist()]

3. 批量获取预测字幕

遍历所有测试图像,调用find_matches获取每个图像的Top1匹配字幕(作为模型预测结果):

predicted_captions = []
# 批量获取所有测试图像的匹配字幕
matches_list = find_matches(text_embeddings, test_image_paths, k=1, normalize=True)
# 提取每个图像的Top1字幕并转为单词列表
for matches in matches_list:
    pred_caption = matches[0].split()
    predicted_captions.append(pred_caption)

4. 计算corpus_bleu分数

在收集完所有真实字幕和预测字幕后,直接调用计算:

# 默认计算BLEU-4(n-gram权重为(0.25,0.25,0.25,0.25)),可自定义权重调整计算BLEU-1/2/3
bleu_score = corpus_bleu(true_captions, predicted_captions)
print(f"Corpus BLEU score: {bleu_score:.4f}")

关键注意事项

  • 原find_matches函数存在变量未定义问题(未从queries中读取image_path),已在上述代码中修正
  • 若每个图像对应多个真实字幕,只需在true_captions的子列表中添加更多分词后的字幕即可

内容的提问来源于stack exchange,提问作者user19686684

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 15:45:41