You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Keras微调数据集后,适配LoRA模型遇TypeError错误求助

解决GemmaTokenizer输入类型不匹配的TypeError

错误根源

你遇到的错误核心是:输入模型的data中文本数据被转换成了float32类型,但GemmaTokenizer要求输入必须是字符串类型,导致SentencepieceTokenizeOp无法处理。调整epochs或batch_size对这个类型不匹配问题没有帮助。

具体解决步骤

  • 回溯数据集预处理流程,排查是否有操作误将字符串文本转成了数值类型(比如错误使用归一化、tf.convert_to_tensor时指定了float dtype)。
  • 重新预处理数据集,强制文本字段为字符串类型:
    import tensorflow as tf
    
    def preprocess_text(examples):
        # 确保文本字段是字符串类型
        examples["text"] = tf.cast(examples["text"], tf.string)
        # 调用tokenizer处理文本
        return tokenizer(
            examples["text"],
            padding="max_length",
            truncation=True,
            max_length=512
        )
    
    # 应用预处理并验证数据类型
    processed_data = raw_dataset.map(preprocess_text)
    # 检查样本类型是否正确
    for batch in processed_data.take(1):
        print("输入文本原始类型:", batch["text"].dtype)  # 应为string
        print("Tokenized后类型:", batch["input_ids"].dtype)  # 应为int32
    
  • 如果使用tf.data.Dataset构建数据集,避免对文本列使用tf.strings.to_number这类会改变类型的函数。
  • 确认LoRA适配后的模型输入层与tokenizer的输出匹配,模型接收的应该是tokenizer生成的int32张量,而非原始的float数据。

快速排查方法

打印数据集样本,直接定位类型错误来源:

sample = next(iter(data))
print("样本内容及类型:", sample)

内容的提问来源于stack exchange,提问作者Sharon

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.17 02:37:09