You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何制作OpenAI CLIP格式图文对及对应代码错误排查

问题根因
  • 循环嵌套逻辑错误:处理单张图片的单条字幕时,嵌套了全量图片文件名遍历,导致每生成一条字幕就会覆盖所有txt文件的内容,最终所有txt仅保留最后一次循环执行的那条字幕,内容完全一致
  • 文件打开模式错误:使用w(覆盖写入)模式打开文件,每次写入都会清空原有内容,即使对应关系正确也只会保留最后1条字幕
  • 图片与字幕绑定逻辑缺失:生成txt文件时没有和当前处理的图片做绑定,完全打乱了图片和字幕的对应关系
修复后可运行代码
import os
from google.colab import drive
drive.mount('/content/drive')

path = '/content/drive/My Drive/Godard_imgs/Eloge_jpgs/eloge_sample/'
images = []

# 收集目标路径下所有图片文件
for filename in os.listdir(path):
    if filename.endswith((".png", ".jpg")):
        images.append(os.path.join(path, filename))

# 逐张处理图片生成对应字幕文件
for img_path in images:
    # 生成当前图片的5条字幕
    captions = caption_image(img_path, args, net, preprocess)
    # 提取当前图片的文件名,生成对应txt文件路径
    img_filename = os.path.basename(img_path)
    txt_file_path = os.path.join(path, f"{img_filename}.txt")
    # 一次性写入所有字幕
    with open(txt_file_path, "w", encoding="utf-8") as f:
        for cap in captions:
            f.write(f"{cap}\n")

内容的提问来源于stack exchange,提问作者Daniel Rasmussen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.26 03:36:04