You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何修改Llama3.2-vision调用脚本,强制输出仅为指定选项列表中的单一值

如何修改Llama3.2-vision调用脚本,强制输出仅为指定选项列表中的单一值

我明白你的困扰——即使prompt写得很明确,大模型偶尔还是会输出多余的解释性内容。要解决这个问题,我们可以从强化prompt约束和添加结果后处理校验两个层面入手,双管齐下确保输出严格符合要求。

具体修改方案

1. 优化Prompt,从源头减少冗余输出

把原来的prompt改成更强硬、无歧义的表述,明确要求模型只能返回指定的分类词,不能有任何额外内容:

"Classify this image into one of these exact categories, and ONLY return the single category word with no extra text, explanation, or punctuation:\n"
"- dog\n"
"- cat\n"
"- butterfly\n"

通过强调ONLY return the single category word with no extra text,最大限度压缩模型的“发挥空间”。

2. 添加结果后处理校验,兜底确保输出合规

即使模型偶尔不听话,我们可以通过代码自动清理返回内容,提取出符合要求的分类词。具体步骤:

  • 定义统一的分类选项常量
  • 编写清理函数,遍历返回内容匹配合法选项
  • 处理大小写不一致的情况(比如模型返回Butterfly转成butterfly)

修改后的完整代码

from pathlib import Path
import base64
import requests

# 统一管理允许的分类选项,方便后续修改
ALLOWED_OPTIONS = ['dog', 'cat', 'butterfly']

def encode_image_to_base64(image_path):
    """Convert an image file to base64 string."""
    return base64.b64encode(image_path.read_bytes()).decode('utf-8')

def clean_response(content, allowed_options):
    """清理模型返回内容,确保仅输出允许的分类词"""
    # 转换为小写,处理大小写不匹配的情况
    content_lower = content.strip().lower()
    # 遍历合法选项,检查是否存在于返回内容中
    for option in allowed_options:
        if option in content_lower:
            return option
    # 未匹配到任何选项时返回默认标记,可根据需求调整
    return 'unknown'

def extract_text_from_image(image_path):
    """Send image to local Llama API and get cleaned classification result."""
    base64_image = encode_image_to_base64(image_path)

    payload = {
        "model": "llama3.2-vision",
        "stream": False,
        "messages": [
            {
                "role": "user",
                "content": (
                    "Classify this image into one of these exact categories, "
                    "and ONLY return the single category word with no extra text, explanation, or punctuation:\n"
                    "- dog\n"
                    "- cat\n"
                    "- butterfly\n"
                ),
                "images": [base64_image]
            }
        ]
    }

    response = requests.post(
        "http://localhost:11434/api/chat",
        json=payload,
        headers={"Content-Type": "application/json"}
    )

    # 获取模型原始返回内容
    raw_content = response.json().get('message', {}).get('content', 'No text extracted')
    # 清理内容,确保输出合规
    cleaned_content = clean_response(raw_content, ALLOWED_OPTIONS)
    return cleaned_content

def process_directory():
    """Process all images in current directory and create text files."""
    for image_path in Path('.').glob('*'):
        if image_path.suffix.lower() in {'.png', '.jpg', '.jpeg', '.gif', '.bmp', '.webp'}:
            print(f"\nProcessing {image_path}...")

            classification_result = extract_text_from_image(image_path)
            output_file = image_path.with_suffix('.txt')
            output_file.write_text(classification_result, encoding='utf-8')
            print(f"Created {output_file} | Classification: {classification_result}")

process_directory()

关键修改点说明

  1. 统一分类选项:用ALLOWED_OPTIONS常量管理所有合法分类,后续修改分类时只需改这一处
  2. 结果清理函数:clean_response作为兜底逻辑,即使模型输出类似"From the image...ANSWER: Butterfly."的冗余内容,也能准确提取出butterfly
  3. 大小写兼容:自动将返回内容转小写,避免因模型返回首字母大写的词导致匹配失败
  4. 透明化日志:在处理目录时打印分类结果,方便你实时校验输出是否符合预期

如果模型完全无法识别图片(返回内容中没有任何合法选项),脚本会返回unknown,你可以根据需求修改这个默认值,或者添加重新调用模型的逻辑。

备注:内容来源于stack exchange,提问作者coolhand

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.13 20:09:32