Qwen2.5-Coder-1.5B SFT训练报错:梯度计算中断问题求助
Qwen2.5-Coder-1.5B SFT训练报错"element 0 of tensors does not require grad and does not have a grad_fn"分析与解决
问题概述
在对Qwen2.5-Coder-1.5B进行监督微调(SFT)时,反向传播阶段触发错误,核心报错信息:
element 0 of tensors does not require grad and does not have a grad_fn
同时日志中出现两个关键提示:
UserWarning: None of the inputs have requires_grad=True. Gradients will be None
No label_names provided for model class `PeftModelForCausalLM`. Since `PeftModel` hides base models input arguments, if label_names is not given, label_names can't be set automatically within `Trainer`. Note that empty label_names list will be used instead.
原始代码
# Download model import os import tokenize from huggingface_hub import snapshot_download from transformers import AutoTokenizer, AutoModelForCausalLM, DataCollatorForLanguageModeling model_id = "Qwen/Qwen2.5-Coder-1.5B" save_dir = f"/root/autodl-tmp/NL2SQL/models/{model_id[5:]}/" os.makedirs(save_dir, exist_ok=True) # snapshot_download(repo_id=model_id, local_dir=save_dir) # Load model model = AutoModelForCausalLM.from_pretrained(save_dir, device_map="cuda") tokenizer = AutoTokenizer.from_pretrained(save_dir, device_map="cuda") # Data processing import pandas as pd from datasets import Dataset # Read CSV file and create Dataset data_dir = '/root/autodl-tmp/NL2SQL/cot-qa.csv' df = pd.read_csv(data_dir) dataset = Dataset.from_pandas(df) def combined_preprocess(batch): texts = [] # Iterate over each sample to construct the complete prompt and completion text for q, a, t in zip(batch["query"], batch["answer"], batch["thinking_process"]): question = str(q) answer = str(a) thinking = str(t) prompt = ( f"For the question: {question}.\n" "Please think step by step, list your thinking process between <think> and </think> and then show the final SQL answer:" ) completion = ( f"<think>{thinking}</think>\nMy final answer is: ```sql\n{answer}\n```" ) texts.append(prompt + "\n" + completion) # Do not perform padding or return torch.Tensor; return a list for the collator to pad later tokenized = tokenizer( texts, truncation=True, max_length=1024 * 2, padding=False, ) return tokenized processed_dataset = dataset.map(combined_preprocess, batched=True, remove_columns=dataset.column_names) # print(processed_dataset[0]) # LoRA configuration from peft import LoraConfig, TaskType, get_peft_model lora_config = LoraConfig( task_type=TaskType.CAUSAL_LM, target_modules=["q_proj", "k_proj", "v_proj", "o_proj", "gate_proj", "up_proj", "down_proj"], r=8, lora_alpha=16, # 8*2 lora_dropout=0.05, bias='none', inference_mode=False ) model = get_peft_model(model, lora_config) print(model.print_trainable_parameters()) model.config.use_cache = False # Training configuration from transformers import TrainingArguments, Trainer training_args = TrainingArguments( output_dir="./output/sft/", per_device_train_batch_size=4, gradient_accumulation_steps=4, logging_steps=10, logging_first_step=5, num_train_epochs=2, save_steps=100, learning_rate=1e-4, save_on_each_node=True, gradient_checkpointing=True, report_to="none", remove_unused_columns=False, ) # Swanlab setup import swanlab from swanlab.integration.transformers import SwanLabCallback swanlab_callback = SwanLabCallback( ... ) trainer = Trainer( model=model, args=training_args, train_dataset=processed_dataset, data_collator=DataCollatorForLanguageModeling(tokenizer=tokenizer, mlm=False), callbacks=[swanlab_callback], ) trainer.train()
完整日志输出
root@autodl-container:~/autodl-tmp/NL2SQL# python sft.py Sliding Window Attention is enabled but not implemented for `sdpa`; unexpected results may be encountered. Map: 100%|██████████████████████████████████████████████████████████████████████████████████████████████████████| 9399/9399 [00:03<00:00, 2396.28 examples/s] trainable params: 9,232,384 || all params: 1,552,946,688 || trainable%: 0.5945 None No label_names provided for model class `PeftModelForCausalLM`. Since `PeftModel` hides base models input arguments, if label_names is not given, label_names can't be set automatically within `Trainer`. Note that empty label_names list will be used instead. swanlab: Tracking run with swanlab version 0.4.11 swanlab: Run data will be saved locally in /root/autodl-tmp/NL2SQL/swanlog/run- swanlab: 👋 Hi , welcome to swanlab! swanlab: Syncing run to the cloud swanlab: 🌟 Run `swanlab watch /root/autodl-tmp/NL2SQL/swanlog` to view SwanLab Experiment Dashboard locally swanlab: 🏠 View project at https://swanlab.cn/@/Qwen2.5-Coder-1.5B-NL2SQL-SFT swanlab: 🚀 View run at https://swanlab.cn/@/Qwen2.5-Coder-1.5B-NL2SQL-SFT/runs/ 0%| | 0/1174 [00:00<?, ?it/s]/root/miniconda3/lib/python3.12/site-packages/torch/utils/checkpoint.py:87: UserWarning: None of the inputs have requires_grad=True. Gradients will be None warnings.warn( swanlab: Error happened while training swanlab: 🌟 Run `swanlab watch /root/autodl-tmp/NL2SQL/swanlog` to view SwanLab Experiment Dashboard locally swanlab: 🏠 View project at https://swanlab.cn/@/Qwen2.5-Coder-1.5B-NL2SQL-SFT swanlab: 🚀 View run at https://swanlab.cn/@/Qwen2.5-Coder-1.5B-NL2SQL-SFT/runs/ File "/root/autodl-tmp/NL2SQL/sft.py", line 117, in <module> trainer.train() File "/root/miniconda3/lib/python3.12/site-packages/transformers/trainer.py", line 2241, in train return inner_training_loop( ^^^^^^^^^^^^^^^^^^^^ File "/root/miniconda3/lib/python3.12/site-packages/transformers/trainer.py", line 2548, in _inner_training_loop tr_loss_step = self.training_step(model, inputs, num_items_in_batch) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/root/miniconda3/lib/python3.12/site-packages/transformers/trainer.py", line 3740, in training_step self.accelerator.backward(loss, **kwargs) File "/root/miniconda3/lib/python3.12/site-packages/accelerate/accelerator.py", line 2329, in backward loss.backward(**kwargs) File "/root/miniconda3/lib/python3.12/site-packages/torch/_tensor.py", line 626, in backward torch.autograd.backward( File "/root/miniconda3/lib/python3.12/site-packages/torch/autograd/__init__.py", line 347, in backward _engine_run_backward( File "/root/miniconda3/lib/python3.12/site-packages/torch/autograd/graph.py", line 823, in _engine_run_backward return Variable._execution_engine.run_backward( # Calls into the C++ engine to run the backward pass ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ element 0 of tensors does not require grad and does not have a grad_fn 0%| | 0/1174 [00:02<?, ?it/s]
诱因分析
- 梯度检查点与LoRA的兼容性问题:开启
gradient_checkpointing=True后,PyTorch梯度检查机制要求输入张量具备requires_grad=True属性,但LoRA封装后的模型默认不会为输入张量启用梯度,导致无法构建有效的梯度计算图,最终触发反向传播错误。 - 缺失label_names参数:Peft模型的封装结构会隐藏基础模型的输入参数,Trainer无法自动识别标签字段,空的label_names列表会导致损失计算与可训练的LoRA权重脱节,梯度无法传递到目标参数。
- 数据处理的潜在问题:若预处理后的数据集未正确生成
input_ids,或DataCollatorForLanguageModeling未成功生成labels字段,会导致损失张量与模型参数无关联,进而无法计算梯度。
解决建议
1. 修复梯度检查点设置
两种可选方案:
- 方案一:关闭梯度检查点:修改
TrainingArguments中的参数:training_args = TrainingArguments( ... gradient_checkpointing=False, ... ) - 方案二:保留梯度检查点并启用输入梯度:在初始化Peft模型后添加一行代码:
model = get_peft_model(model, lora_config) model.enable_input_require_grads() # 添加此行
2. 显式指定label_names
在Trainer初始化时添加label_names参数:
trainer = Trainer( model=model, args=training_args, train_dataset=processed_dataset, data_collator=DataCollatorForLanguageModeling(tokenizer=tokenizer, mlm=False), callbacks=[swanlab_callback], label_names=["labels"] # 显式指定标签字段 )
3. 确保模型处于训练模式
在Trainer初始化前手动设置模型为训练状态:
model.train()
4. 验证数据处理流程
打印预处理后的数据集样本,确认包含input_ids字段:
print(processed_dataset[0])
确保DataCollatorForLanguageModeling正确生成labels(默认会将input_ids复制为labels)。
内容的提问来源于stack exchange,提问作者Wonster
相关产品推荐
相关产品推荐

