You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用LlamaIndex调用Camel-5B模型时触发indexSelectSmallIndex断言失败

问题解决:LlamaIndex使用中触发CUDA索引断言错误

错误信息

/opt/pytorch/pytorch/aten/src/ATen/native/cuda/Indexing.cu:1237:
indexSelectSmallIndex: block: [18,0,0], thread: [31,0,0] Assertion srcIndex < srcSelectDimSize failed.

前置警告提示:

Token indices sequence length is longer than the specified maximum sequence length for this model (1815 > 512). Running this sequence through the model will result in indexing errors
Setting pad_token_id to eos_token_id:50256 for open-end generation.
This is a friendly reminder - the current text generation call will exceed the model's predefined maximum length (2048). Depending on the model, you may observe exceptions, performance degradation, or nothing at all.

问题根源

核心原因是输入token序列长度超出模型限制,结合代码细节,具体问题点:

  • Tokenizer的max_length设置为1024,远小于模型Writer/camel-5b-hf的实际上下文窗口(2048),导致token截断逻辑混乱
  • 重复创建ServiceContext,配置出现冲突
  • 文本块(chunk)大小设置为1024,加上prompt模板和查询语句,总长度容易突破模型的2048上下文限制

修复步骤

1. 统一模型与Tokenizer的长度配置

确保Tokenizer的最大长度和模型上下文窗口保持一致,同时移除重复的ServiceContext创建逻辑:

import torch
from llama_index import VectorStoreIndex, SimpleDirectoryReader, ServiceContext
from llama_index.llms import HuggingFaceLLM
from llama_index.prompts import PromptTemplate

# 加载数据(补充之前缺失的加载步骤)
reader = SimpleDirectoryReader(input_files=["./data/paul_graham/paul_graham_essay1.txt"])
documents = reader.load_data()

# 设置适配的Prompt模板
query_wrapper_prompt = PromptTemplate(
    "Below is an instruction that describes a task. "
    "Write a response that appropriately completes the request.\n\n"
    "### Instruction:\n{query_str}\n\n### Response:"
)

# 加载模型,统一长度配置
llm = HuggingFaceLLM(
    context_window=2048,
    max_new_tokens=256,
    generate_kwargs={"temperature": 0.25, "do_sample": True},
    query_wrapper_prompt=query_wrapper_prompt,
    tokenizer_name="Writer/camel-5b-hf",
    model_name="Writer/camel-5b-hf",
    device_map="auto",
    tokenizer_kwargs={"max_length": 2048},  # 和模型上下文窗口保持一致
    model_kwargs={"torch_dtype": torch.float16}
)

2. 合理设置文本块大小

将chunk_size调整为更小的值(比如512),确保拼接后的prompt+上下文+生成token的总长度不超过2048:

# 仅创建一次ServiceContext
service_context = ServiceContext.from_defaults(chunk_size=512, llm=llm, embed_model="local")

# 建立索引
index = VectorStoreIndex(documents, service_context=service_context)

3. 优化查询时的上下文拼接限制

给查询引擎添加上下文窗口限制,避免总长度溢出:

query_engine = index.as_query_engine(
    service_context=service_context,
    similarity_top_k=3  # 减少召回的文本块数量,控制总长度
)

response = query_engine.query(
    "Whatints可为 ringidd CA_allPrime变形 IncludingBryan probR advance在这里,还要注意什么?"
)
print(response)

额外注意事项

  • 若仍出现内存相关错误,可以尝试降低similarity_top_k的值(比如设为2)
  • 若CUDA显存不足,可启用模型量化减少内存占用:
    model_kwargs={"torch_dtype": torch.float16, "load_in_4bit": True}
    

内容的提问来源于stack exchange,提问作者venergiac

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.01 20:30:40