You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用RoBERTa调用textEmbed函数处理文本时出现尺寸不兼容报错

问题描述

使用text包调用BERT、Electra等模型时运行正常,但调用roberta-base或xlm-roberta-base处理部分文本时频繁报错。

触发报错的文本示例:

627% but if they had lower striked than 16 I would have gone even further OTM. This could really fall off a cliff.

报错信息:

Error in `dplyr::bind_cols()`: ! Can't recycle `..1` (size 44) to match `..2` (size 32)

回溯栈信息:

▆
  1. ├─base::system.time(embeddings_CL2 <- textEmbed(xx, model = "roberta-base"))
  2. ├─text::textEmbed(xx, model = "roberta-base")
  3. │ └─text::textEmbedRawLayers(...)
  4. │   └─text:::sortingLayers(x = hg_embeddings, layers = layers, return_tokens = return_tokens)
  5. │     └─dplyr::bind_cols(tokens_layer_number, layers_4_token)
  6. │       └─vctrs::vec_cbind(!!!dots, .name_repair = .name_repair, .error_call = current_env())
  7. └─vctrs::stop_incompatible_size(...)
  8.   └─vctrs:::stop_incompatible(...)
  9.     └─vctrs:::stop_vctrs(...)
 10.       └─rlang::abort(message, class = c(class, "vctrs_error"), ..., call = call)
报错原因

这个错误的核心是token序列长度不匹配:

  • RoBERTa系列模型的分词器对特殊字符(如示例中的%)、行业缩写(如OTM)的处理逻辑和BERT/Electra存在差异,会生成数量不同的token。
  • 在text包的textEmbed流程中,sortingLayers函数需要将token元数据(tokens_layer_number,长度44)与对应层的嵌入向量(layers_4_token,长度32)按列绑定,但两者长度不一致,触发dplyr::bind_cols的回收校验错误。
  • 本质是text包在处理RoBERTa系列模型时,分词器输出的token数量和模型嵌入层输出的token维度未对齐,属于工具包的兼容性bug,通常和特殊文本片段的分词处理逻辑有关。

内容的提问来源于stack exchange,提问作者Luigi Curini

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 19:42:22