RobertaTokenizer编码结构化Profile句子均得相同值的原因排查
编码异常原因及修正方案
核心问题分析
你的编码逻辑存在多处错误,和句子结构一致无关,具体问题如下:
tokenizer.encode用法错误:tokenizer.encode是处理单个文本字符串的方法,不支持axis参数,你传入的axis=1属于无效参数;同时list(row)会把DataFrame行对象拆成列元素列表,要么是单个字符串被套进列表,要么是零散字段的集合,完全不符合tokenizer的输入要求。apply逻辑错误:默认apply是按列处理数据,你没有指定axis=1,导致整列数据被当成单个输入处理,最终编码出异常值。- 张量处理错误:
squeeze(0)会错误压缩批量数据的维度,破坏输入结构。
修正后的代码示例
假设你的DataFrame有一列profile存储完整的句子文本:
from transformers import RobertaForSequenceClassification, RobertaTokenizer import pandas as pd import torch model_name = "roberta-base" model = RobertaForSequenceClassification.from_pretrained(model_name) tokenizer = RobertaTokenizer.from_pretrained(model_name) def encode_df(dataframe): # 提取所有profile文本,批量编码 texts = dataframe["profile"].tolist() encoding = tokenizer.batch_encode_plus( texts, add_special_tokens=True, padding="longest", truncation=True, return_tensors="pt" ) return encoding["input_ids"] # 读取数据 df = pd.read_excel("/content/profile sentences into df.xlsx") profiles_input = encode_df(df) # 导出ONNX torch.onnx.export( model, profiles_input, ONNX_path, input_names=['input'], output_names=['output'], dynamic_axes={ 'input': {0: 'batch_size', 1: 'sentence_length'}, 'output': {0: 'batch_size'} } )
额外说明
- 之前多列DataFrame的“正常”是巧合:零散字段被错误编码但未出现明显异常值,实际编码逻辑同样不正确。
- 若你是拆分句子为多列存储(对应原fullName、Location等字段),需先拼接成完整句子再编码:
def combine_profile(row): return f"My name is {row['fullName']}. I live in {row['Location']}. I went to {row['College']} in {row['Degree_location']}. I have a degree in {row['Degree']}. I have worked as {row['Jobs']} in companies {row['Company']}. I have worked from {row['Date']}. My job locations have been {row['LocationJob']}. The average distance between my locations is {row['Avg_rounded']}." df["profile"] = df.apply(combine_profile, axis=1)
内容的提问来源于stack exchange,提问作者Jesper Ezra
相关产品推荐
相关产品推荐

