使用LlamaIndex调用Camel-5B模型时触发indexSelectSmallIndex断言失败
错误信息
/opt/pytorch/pytorch/aten/src/ATen/native/cuda/Indexing.cu:1237:
indexSelectSmallIndex: block: [18,0,0], thread: [31,0,0] AssertionsrcIndex < srcSelectDimSizefailed.
前置警告提示:
Token indices sequence length is longer than the specified maximum sequence length for this model (1815 > 512). Running this sequence through the model will result in indexing errors
Settingpad_token_idtoeos_token_id:50256 for open-end generation.
This is a friendly reminder - the current text generation call will exceed the model's predefined maximum length (2048). Depending on the model, you may observe exceptions, performance degradation, or nothing at all.
问题根源
核心原因是输入token序列长度超出模型限制,结合代码细节,具体问题点:
- Tokenizer的
max_length设置为1024,远小于模型Writer/camel-5b-hf的实际上下文窗口(2048),导致token截断逻辑混乱 - 重复创建
ServiceContext,配置出现冲突 - 文本块(chunk)大小设置为1024,加上prompt模板和查询语句,总长度容易突破模型的2048上下文限制
修复步骤
1. 统一模型与Tokenizer的长度配置
确保Tokenizer的最大长度和模型上下文窗口保持一致,同时移除重复的ServiceContext创建逻辑:
import torch from llama_index import VectorStoreIndex, SimpleDirectoryReader, ServiceContext from llama_index.llms import HuggingFaceLLM from llama_index.prompts import PromptTemplate # 加载数据(补充之前缺失的加载步骤) reader = SimpleDirectoryReader(input_files=["./data/paul_graham/paul_graham_essay1.txt"]) documents = reader.load_data() # 设置适配的Prompt模板 query_wrapper_prompt = PromptTemplate( "Below is an instruction that describes a task. " "Write a response that appropriately completes the request.\n\n" "### Instruction:\n{query_str}\n\n### Response:" ) # 加载模型,统一长度配置 llm = HuggingFaceLLM( context_window=2048, max_new_tokens=256, generate_kwargs={"temperature": 0.25, "do_sample": True}, query_wrapper_prompt=query_wrapper_prompt, tokenizer_name="Writer/camel-5b-hf", model_name="Writer/camel-5b-hf", device_map="auto", tokenizer_kwargs={"max_length": 2048}, # 和模型上下文窗口保持一致 model_kwargs={"torch_dtype": torch.float16} )
2. 合理设置文本块大小
将chunk_size调整为更小的值(比如512),确保拼接后的prompt+上下文+生成token的总长度不超过2048:
# 仅创建一次ServiceContext service_context = ServiceContext.from_defaults(chunk_size=512, llm=llm, embed_model="local") # 建立索引 index = VectorStoreIndex(documents, service_context=service_context)
3. 优化查询时的上下文拼接限制
给查询引擎添加上下文窗口限制,避免总长度溢出:
query_engine = index.as_query_engine( service_context=service_context, similarity_top_k=3 # 减少召回的文本块数量,控制总长度 ) response = query_engine.query( "Whatints可为 ringidd CA_allPrime变形 IncludingBryan probR advance在这里,还要注意什么?" ) print(response)
额外注意事项
- 若仍出现内存相关错误,可以尝试降低
similarity_top_k的值(比如设为2) - 若CUDA显存不足,可启用模型量化减少内存占用:
model_kwargs={"torch_dtype": torch.float16, "load_in_4bit": True}
内容的提问来源于stack exchange,提问作者venergiac

