You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何让本地Llama 2调用RTX 3060 Ti GPU而非CPU?

问题解决指导

一、GPU加速配置(Windows环境)

你的llama-cpp-python默认是CPU版本,需重新安装带CUDA支持的版本来启用GPU加速,步骤如下:

  1. 卸载现有版本:
    pip uninstall -y llama-cpp-python
    
  2. 安装带CUDA加速的版本(提前确保已安装NVIDIA CUDA Toolkit 11.8/12.x,适配RTX3060Ti):
    set CMAKE_ARGS="-DLLAMA_CUBLAS=on"
    set FORCE_CMAKE=1
    pip install llama-cpp-python --force-reinstall --upgrade --no-cache-dir
    
  3. 加载模型时启用GPU加速,修改初始化代码:
    LLM = Llama(
        model_path=r"D:\VoiceAssisant\llama-2-13b-chat.Q6_K.gguf",
        f16_kv=True,
        n_gpu_layers=-1,  # 内存足够时将所有模型层加载到GPU;显存不足可设为40-60
        n_ctx=4096  # 增大上下文窗口,适配长对话
    )
    

二、回答异常问题修复

1. 回答截断问题

当前代码max_tokens=256限制了输出长度,直接调大即可:

output = LLM(prompt, max_tokens=1024, stop=["</s>"], echo=False)

2. 回答偏离主题问题

Llama 2 Chat模型需使用官方指定的聊天格式,替换原有的Q:A模板:

prompt = f"<s>[INST] {input_text} [/INST]"

同时将stop参数改为模型默认停止词["</s>"],避免提前截断。

三、修改后的完整代码

import speech_recognition as sr
import pyttsx3
import warnings
from llama_cpp import Llama

warnings.filterwarnings("ignore")

# 初始化语音识别器
recognizer = sr.Recognizer()

# 初始化文字转语音引擎
engine = pyttsx3.init()

# 唤醒词和休眠词设置
wake_up_word = "assistant"
sleep_mode_word = "mute"

# 加载模型并启用GPU加速
LLM = Llama(
    model_path=r"D:\VoiceAssisant\llama-2-13b-chat.Q6_K.gguf",
    f16_kv=True,
    n_gpu_layers=-1,
    n_ctx=4096
)

def listen():
    asleep = True

    while True:
        print("Listening for wake up word...")
        with sr.Microphone() as source:
            audio = recognizer.listen(source)
            try:
                text = recognizer.recognize_google(audio)
                if text.lower() == wake_up_word:
                    print("Listening for query...")
                    asleep = False
                    while True:
                        with sr.Microphone() as source:
                            audio = recognizer.listen(source)
                            try:
                                input_text = recognizer.recognize_google(audio)
                                if input_text.lower() == sleep_mode_word:
                                    print("Sleep mode activated...")
                                    asleep = True
                                    break
                                else:
                                    # 使用Llama 2官方聊天格式
                                    prompt = f"<s>[INST] {input_text} [/INST]"
                                    output = LLM(prompt, max_tokens=1024, stop=["</s>"], echo=False)
                                    response_text = output["choices"][0]["text"].strip()
                                    print(f"Response: {response_text}")
                                    engine.say(response_text)
                                    engine.runAndWait()
                            except sr.UnknownValueError:
                                print("无法识别语音内容")
                elif text.lower() == sleep_mode_word:
                    print("Sleep mode activated...")
                    asleep = True
            except sr.UnknownValueError:
                pass

listen()

内容的提问来源于stack exchange,提问作者AntoineCra

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.09 08:38:12