如何修改Llama3.2-vision调用脚本,强制输出仅为指定选项列表中的单一值
如何修改Llama3.2-vision调用脚本,强制输出仅为指定选项列表中的单一值
我明白你的困扰——即使prompt写得很明确,大模型偶尔还是会输出多余的解释性内容。要解决这个问题,我们可以从强化prompt约束和添加结果后处理校验两个层面入手,双管齐下确保输出严格符合要求。
具体修改方案
1. 优化Prompt,从源头减少冗余输出
把原来的prompt改成更强硬、无歧义的表述,明确要求模型只能返回指定的分类词,不能有任何额外内容:
"Classify this image into one of these exact categories, and ONLY return the single category word with no extra text, explanation, or punctuation:\n" "- dog\n" "- cat\n" "- butterfly\n"
通过强调ONLY return the single category word with no extra text,最大限度压缩模型的“发挥空间”。
2. 添加结果后处理校验,兜底确保输出合规
即使模型偶尔不听话,我们可以通过代码自动清理返回内容,提取出符合要求的分类词。具体步骤:
- 定义统一的分类选项常量
- 编写清理函数,遍历返回内容匹配合法选项
- 处理大小写不一致的情况(比如模型返回
Butterfly转成butterfly)
修改后的完整代码
from pathlib import Path import base64 import requests # 统一管理允许的分类选项,方便后续修改 ALLOWED_OPTIONS = ['dog', 'cat', 'butterfly'] def encode_image_to_base64(image_path): """Convert an image file to base64 string.""" return base64.b64encode(image_path.read_bytes()).decode('utf-8') def clean_response(content, allowed_options): """清理模型返回内容,确保仅输出允许的分类词""" # 转换为小写,处理大小写不匹配的情况 content_lower = content.strip().lower() # 遍历合法选项,检查是否存在于返回内容中 for option in allowed_options: if option in content_lower: return option # 未匹配到任何选项时返回默认标记,可根据需求调整 return 'unknown' def extract_text_from_image(image_path): """Send image to local Llama API and get cleaned classification result.""" base64_image = encode_image_to_base64(image_path) payload = { "model": "llama3.2-vision", "stream": False, "messages": [ { "role": "user", "content": ( "Classify this image into one of these exact categories, " "and ONLY return the single category word with no extra text, explanation, or punctuation:\n" "- dog\n" "- cat\n" "- butterfly\n" ), "images": [base64_image] } ] } response = requests.post( "http://localhost:11434/api/chat", json=payload, headers={"Content-Type": "application/json"} ) # 获取模型原始返回内容 raw_content = response.json().get('message', {}).get('content', 'No text extracted') # 清理内容,确保输出合规 cleaned_content = clean_response(raw_content, ALLOWED_OPTIONS) return cleaned_content def process_directory(): """Process all images in current directory and create text files.""" for image_path in Path('.').glob('*'): if image_path.suffix.lower() in {'.png', '.jpg', '.jpeg', '.gif', '.bmp', '.webp'}: print(f"\nProcessing {image_path}...") classification_result = extract_text_from_image(image_path) output_file = image_path.with_suffix('.txt') output_file.write_text(classification_result, encoding='utf-8') print(f"Created {output_file} | Classification: {classification_result}") process_directory()
关键修改点说明
- 统一分类选项:用
ALLOWED_OPTIONS常量管理所有合法分类,后续修改分类时只需改这一处 - 结果清理函数:
clean_response作为兜底逻辑,即使模型输出类似"From the image...ANSWER: Butterfly."的冗余内容,也能准确提取出butterfly - 大小写兼容:自动将返回内容转小写,避免因模型返回首字母大写的词导致匹配失败
- 透明化日志:在处理目录时打印分类结果,方便你实时校验输出是否符合预期
如果模型完全无法识别图片(返回内容中没有任何合法选项),脚本会返回unknown,你可以根据需求修改这个默认值,或者添加重新调用模型的逻辑。
备注:内容来源于stack exchange,提问作者coolhand
相关产品推荐
相关产品推荐

