You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Intel Xeon规格微调Llama3-8B效果不佳,求无RAG优化方案

Llama3-8B微调Intel Xeon参数优化方案(无RAG)

问题描述

使用PEFT框架微调Llama3-8B模型学习Intel Xeon系列CPU的核心数、GFLOPS、缓存、主频等规格参数,完成1-20轮微调后损失从1.1降至0.35,但模型无法输出精确的参数值,寻求无需RAG技术的优化建议。

核心优化建议

一、数据集优化

  • 统一输出格式并强化结构化:当前输出格式冗余(重复CPU名称),改成简洁的结构化格式,让模型聚焦参数本身:
    gigaflops:1094.4, core_count:38, base_frequency:2.40 GHz, cache_size:57M
    
    可加入少量JSON格式样本,强化模型对结构化输出的认知:
    {"gigaflops": 1094.4, "core_count": 38, "base_frequency": "2.40 GHz", "cache_size": "57M"}
    
  • 提升数据覆盖与区分度:确保数据集覆盖全系列Xeon型号,包括不同代际、定位的产品;加入相似型号的对比样本,强化模型对型号差异的识别能力;手动构造难区分型号的样本,提升精准度。
  • 清洗数据一致性:检查并修正数据集中的参数冲突(同一型号不同参数),统一参数表述格式(比如主频保留两位小数、缓存单位统一为M/G),避免模型学习矛盾信息。

二、训练参数调整

  • 优化学习率与训练轮次:将学习率降至5e-61e-5区间,同时增加训练轮次至3050轮,确保模型充分学习参数映射关系。
  • 启用评估与早停:恢复训练集拆分逻辑,保留10%~15%的验证集,开启evaluation_strategy="epoch"和early_stopping_patience=3,避免过拟合并监控验证集性能。
  • 调整批次与梯度累积:将gradient_accumulation_steps提升至2~4,在不增加显存压力的前提下增大有效批次,让参数更新更稳定。
  • 调整NeFTune噪声:将neftune_noise_alpha降至1或直接关闭,避免过多噪声干扰模型对精确数值的学习。

三、PEFT配置优化

  • 扩大LoRA目标层范围:指定更多关键层,增强LoRA的影响能力:
    loracofig_instance = LoraConfig(
        task_type="CAUSAL_LM",
        inference_mode=False,
        r=16,
        use_rslora=True,
        lora_alpha=32,
        lora_dropout=0.05,
        target_modules=["q_proj", "v_proj", "k_proj", "o_proj", "gate_proj", "up_proj", "down_proj"]
    )
    
  • 调整LoRA参数容量:将r提升至32,lora_alpha对应调整为64,增大LoRA参数容量;可尝试关闭use_rslora,用标准LoRA训练对比效果。

四、训练策略优化

  • 关闭样本拼接:将packing=True改为packing=False,让每个样本独立训练,强化模型对单一样本的精准输出能力。
  • 缩小序列长度:将max_seq_length降至512或256,减少无效padding,提升训练效率与模型聚焦能力。
  • 简化提示词模板:统一system提示词,明确任务要求,减少冗余信息:
    <|begin_of_text|><|start_header_id|>system<|end_header_id|>
    给定Intel Xeon处理器型号,输出其GFLOPS、核心数、基础主频、缓存大小,格式为:gigaflops:数值, core_count:数值, base_frequency:数值 GHz, cache_size:数值M
    <|eot_id|>
    <|start_header_id|>user<|end_header_id|>
    Intel Xeon Platinum 8368 Processor<|eot_id|>
    <|start_header_id|>assistant<|end_header_id|>
    gigaflops:1094.4, core_count:38, base_frequency:2.40 GHz, cache_size:57M<|eot_id|>
    

五、推理阶段优化

  • 设置严格解码参数:推理时使用temperature=0、top_k=1、top_p=0.0,强制模型输出最可能的结果;设置max_new_tokens为固定值(比如100),避免冗余输出。
  • 添加格式约束提示:在推理的用户提示词末尾加上格式要求,引导模型输出符合规范的内容。

数据集示例

"<|begin_of_text|><|start_header_id|>system<|end_header_id|> 
 Intel processors names are given as input, your task is to give output of Giga Flops(floating point operations), Core count, base Frequency, and Cache for the input processor.
The first two digits of the CPU name represents processor series(8-platinum, 6/5-Gold, 4-Silver, 3-Bronze) and processor generation  2nd,3rd 4th respectively.
 <|eot_id|> 

 <|start_header_id|>user<|end_header_id|> 
Intel Xeon Platinum 8368 Processor <|eot_id|> 

 <|start_header_id|>assistant<|end_header_id|> 
 gigaflops for Intel Xeon Platinum 8368 Processor:1094.4, core_count for Intel Xeon Platinum 8368 Processor:38, base_frequency for Intel Xeon Platinum 8368 Processor:2.40 GHz, cache_size for Intel Xeon Platinum 8368 Processor:57M  <|eot_id|>."

微调代码

import torch
import pandas as pd
from trl import SFTTrainer
from sklearn.model_selection import train_test_split
from datasets import load_dataset, Dataset
from peft import (
    LoraConfig,
    get_peft_model)

from transformers import (
    AutoModelForCausalLM,
    AutoTokenizer,
    TrainingArguments,
    BitsAndBytesConfig,
    DefaultDataCollator,

)
from torch.utils.data import DataLoader
import torch.nn as nn
import os

df = pd.read_csv("/sda/ak_cpu_prompt.csv")

# train_dataset, test_dataset = train_test_split(df,shuffle=True,test_size=0.1,random_state=42)
# training_data = Dataset.from_pandas(df=train_dataset)
# testing_data = Dataset.from_pandas(df=test_dataset)

training_data = Dataset.from_pandas(df=df)

training_data.set_format(
    type="torch",
    columns=["output"],
    device="cuda",
)
pretrained_model = "/sda/llama3/Meta-Llama-3-8B-Instruct-hf"

bnb_config = BitsAndBytesConfig(
    load_in_4bit=True,
    bnb_4bit_quant_type="nf4",
    bnb_4bit_compute_dtype=torch.bfloat16,
    bnb_4bit_use_double_quant=True,
)

print("\n....Loading model...")
model = AutoModelForCausalLM.from_pretrained(
    pretrained_model,
    quantization_config=bnb_config, ###
device_map='auto')

# model.config.pretrainig_tp = 1  # For Faster responses, trading off accuracy
print("....Loading model complete...\n")

tokenizer = AutoTokenizer.from_pretrained(pretrained_model,padding_side='left')
tokenizer.pad_token = tokenizer.eos_token
# tokenizer.padding_side = "right"
print("...preparing Training arguments...")

#tokenizer.add_special_tokens({'pad_token': '[PAD]'})
training_arg_instance = TrainingArguments(
    output_dir="/sda/llama3/llama_3_adapters/cpu_model-0.1", # change here
    per_device_train_batch_size=4,
    gradient_accumulation_steps=1,
    prediction_loss_only=True,
    gradient_checkpointing=False,
    # evaluation_strategy="steps", # no evaluation dataset
    do_eval=False, # no evaluation dataset
    max_grad_norm=0.3,
    fp16=True,
    bf16=False,
    optim="paged_adamw_32bit",
    logging_steps=10,
    learning_rate=2e-5,
    warmup_ratio=0.0,
    lr_scheduler_type="cosine",
    num_train_epochs=1, # changing for every model
    save_strategy="epoch",
)

print("...preparing Training arguments. Done....\n")


print("...preparing lora config...")
loracofig_instance = LoraConfig(
    task_type="CAUSAL_LM",
    inference_mode=False,
    r=16,
    use_rslora=True,
    lora_alpha=32,
    #lora_dropout=0.01,
    lora_dropout=0.05
)

print("...preparing lora config. Done...\n")

print("...intializing SFT trainer...")
model = get_peft_model(model, loracofig_instance)
#device = [0,1]
#model = get_peft_model(model, loracofig_instance).to(device)

SFT_instance = SFTTrainer(
    model=model,
    train_dataset=training_data,
    # eval_dataset=testing_data,
    dataset_text_field="output",
    max_seq_length=2048,
    tokenizer=tokenizer,
    args=training_arg_instance,
    dataset_batch_size=1,
    neftune_noise_alpha=3,
    # data_collator=DefaultDataCollator,
    packing=True,
    #peft_config=loracofig_instance, # passed in the model 
)
print("...intializing SFT trainer. Done...\n")

model.resize_token_embeddings(len(tokenizer))

trainer = SFT_instance
print("\nStarted training...\n")
trainer.train()
trainer.save_model()
print("\nTraining Comleted, Model saved at : /sda/llama3/llama_3_adapters/cpu_model-0.1") # change here 

内容的提问来源于stack exchange,提问作者AKSHAY JAIN

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 08:44:56