You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于LoRa微调HuggingFace LLM时遭遇CUDA检测失败问题求助

解决CUDA检测失败及bitsandbytes GPU支持缺失问题

问题背景

本地硬件配置为Ryzen 3960x、RTX 3090、64GB内存,测试基于LoRa微调HuggingFace大语言模型,选用《大卫·科波菲尔》训练GPT-2,已完成PDF文本提取、清洗与分词,但微调时触发CUDA检测失败,且bitsandbytes无GPU支持的报错。

报错信息

bin C:\Users\salom\AppData\Local\Packages\PythonSoftwareFoundation.Python.3.11_qbz5n2kfra8p0\LocalCache\local-packages\Python311\site-packages\bitsandbytes\libbitsandbytes_cpu.so
False
C:\Users\salom\AppData\Local\Packages\PythonSoftwareFoundation.Python.3.11_qbz5n2kfra8p0\LocalCache\local-packages\Python311\site-packages\bitsandbytes\cextension.py:34: UserWarning: The installed version of bitsandbytes was compiled without GPU support. 8-bit optimizers, 8-bit multiplication, and GPU quantization are unavailable.
  warn("The installed version of bitsandbytes was compiled without GPU support. "
'NoneType' object has no attribute 'cadam32bit_grad_fp32'
CUDA SETUP: Required library version not found: libbitsandbytes_cpu.so. Maybe you need to compile it from source?
CUDA SETUP: Defaulting to libbitsandbytes_cpu.so...

================================================ERROR=====================================
CUDA SETUP: CUDA detection failed! Possible reasons:
1. CUDA driver not installed
2. CUDA not installed
3. You have multiple conflicting CUDA libraries
4. Required library not pre-compiled for this bitsandbytes release!
CUDA SETUP: If you compiled from source, try again with `make CUDA_VERSION=DETECTED_CUDA_VERSION` for example, `make CUDA_VERSION=113`.
CUDA SETUP: The CUDA version for the compile might depend on your conda install. Inspect CUDA version via `conda list | grep cuda`.
================================================================================

CUDA SETUP: Problem: The main issue seems to be that the main CUDA library was not detected.
CUDA SETUP: Solution 1): Your paths are probably not up-to-date. You can update them via: sudo ldconfig.
CUDA SETUP: Solution 2): If you do not have sudo rights, you can do the following:
CUDA SETUP: Solution 2a): Find the cuda library via: find / -name libcuda.so 2>/dev/null
CUDA SETUP: Solution 2b): Once the library is found add it to the LD_LIBRARY_PATH: export LD_LIBRARY_PATH=$LD_LIBRARY_PATH:FOUND_PATH_FROM_2a
CUDA SETUP: Solution 2c): For a permanent solution add the export from 2b into your .bashrc file, located at ~/.bashrc
CUDA SETUP: Setup Failed!

用户代码

import PyPDF2

# Function to extract text from a PDF file
def extract_text_from_pdf(file_path):
    with open(file_path, 'rb') as file:
        pdf_reader = PyPDF2.PdfReader(file)
        text = ""
        for page in pdf_reader.pages:
            text += page.extract_text()
        return text

# Load the PDF file and extract text
pdf_file_path = "DavidCopperfield.pdf"
book_text = extract_text_from_pdf(pdf_file_path)

import re

# Function to filter and clean the text
def filter_text(text):
    # Remove chapter titles and page numbers
    text = re.sub(r'CHAPTER \d+', '', text)
    text = re.sub(r'\d+', '', text)

    # Remove unwanted characters and extra whitespaces
    text = re.sub(r'[^\w\s\'.-]', '', text)
    text = re.sub(r'\s+', ' ', text)

    # Remove lines with all uppercase letters (potential noise)
    text = '\n'.join(line for line in text.split('\n') if not line.isupper())

    return text

# Apply text filtering to the book text
filtered_text = filter_text(book_text)

# Partition the filtered text into training texts with a maximum size
max_text_size = 150
train_texts = []
current_text = ""
for paragraph in filtered_text.split("\n\n"):
    if len(current_text) + len(paragraph) < max_text_size:
        current_text += paragraph + "\n\n"
    else:
        train_texts.append(current_text)
        current_text = paragraph + "\n\n"
if current_text:
    train_texts.append(current_text)


from transformers import GPT2LMHeadModel, GPT2Tokenizer, GPT2Config
from transformers import AdamW
from torch.utils.data import Dataset, DataLoader
import torch
# Define your dataset class
class TextDataset(Dataset):
    def __init__(self, texts, tokenizer, max_length):
        self.texts = [text for text in texts if len(text) >= max_length]  # Filter out texts shorter than max_length
        self.tokenizer = tokenizer
        self.max_length = max_length

    def __len__(self):
        return len(self.texts)

    def __getitem__(self, idx):
        text = self.texts[idx]
        encoded_input = self.tokenizer.encode_plus(text, max_length=self.max_length, padding='max_length', truncation=True, return_tensors='pt')
        input_ids = encoded_input['input_ids'].squeeze()
        attention_mask = encoded_input['attention_mask'].squeeze()
        return input_ids, attention_mask

# Load pre-trained LM and tokenizer
lm_model = GPT2LMHeadModel.from_pretrained('gpt2')
tokenizer = GPT2Tokenizer.from_pretrained('gpt2')
tokenizer.add_special_tokens({'pad_token': '[PAD]'})  # Add padding token

# Prepare your training data
train_dataset = TextDataset(train_texts, tokenizer, max_length=128)
train_dataloader = DataLoader(train_dataset, batch_size=8, shuffle=True)

# Configure LM training
lm_model.train()
# Replace the optimizer initialization line
optimizer = torch.optim.AdamW(lm_model.parameters(), lr=1e-5)
num_epochs = 10

# Training loop
for epoch in range(num_epochs):
    for batch in train_dataloader:
        input_ids, attention_mask = batch
        outputs = lm_model(input_ids=input_ids, attention_mask=attention_mask, labels=input_ids)
        loss = outputs.loss

        # Backpropagation and optimization
        optimizer.zero_grad()
        loss.backward()
        optimizer.step()

        # Print loss or other metrics for monitoring

# Save the fine-tuned LM
lm_model.save_pretrained('fine_tuned_lm')
tokenizer.save_pretrained('fine_tuned_lm')

解决方案

1. 验证并修复CUDA环境配置

  • 打开命令行执行nvidia-smi,查看驱动支持的CUDA版本(RTX3090适配CUDA 11.3~11.7)
  • 安装对应版本的CUDA Toolkit,确保系统环境变量中添加:
    • CUDA_PATH:指向CUDA安装目录(如C:\Program Files\NVIDIA GPU Computing Toolkit\CUDA\v11.7)
    • PATH:添加%CUDA_PATH%\bin和%CUDA_PATH%\libnvvp
  • 重启终端后执行nvcc --version,确认CUDA版本与nvidia-smi显示的兼容

2. 修复bitsandbytes的GPU支持

  • 卸载当前CPU版本的bitsandbytes:
    pip uninstall bitsandbytes -y
    
  • 安装Windows适配的GPU版本:
    pip install bitsandbytes-windows
    
  • 验证GPU支持:
    python -c "import bitsandbytes; print(bitsandbytes.cuda.is_available())"
    
    返回True则说明GPU支持正常

3. 修改代码以启用GPU加速

在用户代码中添加以下关键修改:

  • 加载模型后将其移至GPU:
    device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
    lm_model = lm_model.to(device)
    
  • 训练循环中把批次数据移至GPU:
    for batch in train_dataloader:
        input_ids, attention_mask = batch
        # 新增:移至GPU
        input_ids = input_ids.to(device)
        attention_mask = attention_mask.to(device)
        outputs = lm_model(input_ids=input_ids, attention_mask=attention_mask, labels=input_ids)
        # ... 后续代码不变
    

4. 处理Windows与Linux路径差异

报错中的sudo ldconfig、LD_LIBRARY_PATH等为Linux环境命令,Windows下无需执行,只需确保CUDA环境变量配置正确即可。


内容的提问来源于stack exchange,提问作者user21537823

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.16 19:34:50