You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Llama 3实现二进制转十进制返回错误结果,求排查帮助

排查Llama 3调用时二进制转十进制返回无意义串的问题

问题场景

在Jupyter Notebook中调用Meta-Llama-3-8B实现二进制转十进制功能,输入二进制串(如"1011")后,模型未返回正确的十进制结果,反而输出无意义的二进制串(如1101 1110...直至达到max_new_tokens上限)。相关代码如下:

!pip install -r requirements.txt
import json
import torch
from transformers import (AutoTokenizer,
                          AutoModelForCausalLM,
                          BitsAndBytesConfig,
                          pipeline)
config_data = json.load(open("config.json"))
HF_TOKEN = config_data["HF_TOKEN"]
model_name = "meta-llama/Meta-Llama-3-8B"
bnb_config = BitsAndBytesConfig(
    load_in_4bit=True,
    bnb_4bit_use_double_quant=True,
    bnb_4bit_quant_type="nf4",
    bnb_4bit_compute_dtype=torch.bfloat16
)
tokenizer = AutoTokenizer.from_pretrained(model_name, use_auth_token=HF_TOKEN)
tokenizer.pad_token = tokenizer.eos_token
model = AutoModelForCausalLM.from_pretrained(
    model_name,
    quantization_config=bnb_config,
    use_auth_token=HF_TOKEN
)
text_generator = pipeline(
    "text-generation",
    model=model,
    tokenizer=tokenizer,
)
def convert_to_decimal(html_input):
    prompt = (f"I'm giving you a input as a binary number"
              f"I need you to convert the binary number to its decimal equivalent. "
              f"Do not give any explanation or add any further text. "
              f"Only provide the transformed output once. "
              f"Don't add any other text except the output. "
              f"For input '101', the output is '5'.\n\n"
              f"{html_input}")
    
    response = text_generator(prompt, max_new_tokens=20, temperature=0.2, return_full_text=False)
    generated_text = response[0]['generated_text'].strip()
    return generated_text
html_input = "1011"
output = convert_to_decimal(html_input)
print(output)

问题原因及修复方案

1. 提示词格式混乱,模型无法正确理解任务

原提示词存在两个关键问题:

  • 第一句"I'm giving you a input as a binary number"末尾无空格/标点,与下一句直接拼接,导致语义混乱;
  • 输入与示例、指令的边界模糊,模型无法识别{html_input}是需要处理的目标,误以为要继续生成二进制内容。

修复后的提示词:

prompt = (f"Convert the following binary number to its decimal equivalent. "
          f"Only output the decimal number, no extra explanation or text. "
          f"Example: Input '101' → Output '5'\n\n"
          f"Input: {html_input}\nOutput:")

通过明确的Input:/Output:标记,清晰划分任务指令、示例和待处理输入,引导模型聚焦于生成十进制结果。

2. 生成参数未限制随机性,导致模型发散

原代码中temperature=0.2虽降低了随机性,但仍可能导致模型偏离任务;未设置do_sample=False时,模型仍会基于概率生成内容,容易出现无意义输出。

调整生成参数:

response = text_generator(
    prompt,
    max_new_tokens=5,  # 十进制结果长度远小于20,缩小上限避免冗余生成
    temperature=0.0,   # 完全消除随机性,强制确定性输出
    do_sample=False,
    pad_token_id=tokenizer.eos_token_id,  # 避免生成pad token
    return_full_text=False
)
  • 将temperature设为0,配合do_sample=False,确保模型严格遵循指令生成结果;
  • 缩小max_new_tokens至合理范围(如5),避免模型生成多余内容。

3. (可选)添加输出格式约束

若模型仍有偏差,可在提示词中明确要求输出为纯数字,例如在示例后补充:"Your output must be a single integer, no other characters allowed."

修复后的完整代码

!pip install -r requirements.txt
import json
import torch
from transformers import (AutoTokenizer,
                          AutoModelForCausalLM,
                          BitsAndBytesConfig,
                          pipeline)
config_data = json.load(open("config.json"))
HF_TOKEN = config_data["HF_TOKEN"]
model_name = "meta-llama/Meta-Llama-3-8B"
bnb_config = BitsAndBytesConfig(
    load_in_4bit=True,
    bnb_4bit_use_double_quant=True,
    bnb_4bit_quant_type="nf4",
    bnb_4bit_compute_dtype=torch.bfloat16
)
tokenizer = AutoTokenizer.from_pretrained(model_name, use_auth_token=HF_TOKEN)
tokenizer.pad_token = tokenizer.eos_token
model = AutoModelForCausalLM.from_pretrained(
    model_name,
    quantization_config=bnb_config,
    use_auth_token=HF_TOKEN
)
text_generator = pipeline(
    "text-generation",
    model=model,
    tokenizer=tokenizer,
)
def convert_to_decimal(html_input):
    prompt = (f"Convert the following binary number to its decimal equivalent. "
              f"Only output the decimal number, no extra explanation or text. "
              f"Example: Input '101' → Output '5'\n\n"
              f"Input: {html_input}\nOutput:")
    
    response = text_generator(
        prompt,
        max_new_tokens=5,
        temperature=0.0,
        do_sample=False,
        pad_token_id=tokenizer.eos_token_id,
        return_full_text=False
    )
    generated_text = response[0]['generated_text'].strip()
    return generated_text
html_input = "1011"
output = convert_to_decimal(html_input)
print(output)  # 预期输出:11

内容的提问来源于stack exchange,提问作者L lawliet

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.23 08:25:21