使用HuggingFace模型时遇VLLM最大序列长度错误
VLLM调用时KV缓存与序列长度不匹配问题解决
报错信息
lm = VLLM( File "/home/ubuntu/isolated_product_description/ipd/lib/python3.8/site-packages/langchain_core/load/serializable.py", line 120, in __init__ super().__init__(**kwargs) File "/home/ubuntu/isolated_product_description/ipd/lib/python3.8/site-packages/pydantic/v1/main.py", line 341, in __init__ raise validation_error pydantic.v1.error_wrappers.ValidationError: 1 validation error for VLLM __root__ The model's max seq len (32768) is larger than the maximum number of tokens that can be stored in KV cache (32624). Try increasing `gpu_memory_utilization` or decreasing `max_model_len` when initializing the engine. (type=value_error)
尝试过的代码
from langchain_community.llms import VLLM llm = VLLM( vllm_kwargs={"quantization": "awq"}, max_model_len=30624, model="TheBloke/Mistral-7B-Instruct-v0.2-AWQ", # gpu_memory_utilization=1.0, trust_remote_code=True, # mandatory for hf models max_new_tokens=512, speculative_max_model_len = 30624, top_k=40, top_p=0.95, temperature=0.7, repetition_penalty= 1.1, )
解决方法
问题核心是参数传递方式错误:LangChain的VLLM类中,max_model_len需要放入vllm_kwargs字典才能被底层vLLM引擎识别,顶层参数不会生效。
修改后的代码:
from langchain_community.llms import VLLM llm = VLLM( vllm_kwargs={ "quantization": "awq", "max_model_len": 30624, "gpu_memory_utilization": 0.95 # 可选:提高GPU内存利用率,扩大KV缓存容量 }, model="TheBloke/Mistral-7B-Instruct-v0.2-AWQ", trust_remote_code=True, max_new_tokens=512, top_k=40, top_p=0.95, temperature=0.7, repetition_penalty=1.1, )
关键说明
- 将
max_model_len移至vllm_kwargs内,确保底层vLLM引擎读取到该配置,让模型最大序列长度匹配KV缓存容量。 - 若GPU内存充足,可将
gpu_memory_utilization设为0.95或1.0,进一步扩大KV缓存的可用空间,避免后续再出现类似问题。
内容的提问来源于stack exchange,提问作者Eric
相关产品推荐
相关产品推荐

