You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

RobertaTokenizer编码结构化Profile句子均得相同值的原因排查

编码异常原因及修正方案

核心问题分析

你的编码逻辑存在多处错误,和句子结构一致无关,具体问题如下:

  • tokenizer.encode用法错误:tokenizer.encode是处理单个文本字符串的方法,不支持axis参数,你传入的axis=1属于无效参数;同时list(row)会把DataFrame行对象拆成列元素列表,要么是单个字符串被套进列表,要么是零散字段的集合,完全不符合tokenizer的输入要求。
  • apply逻辑错误:默认apply是按列处理数据,你没有指定axis=1,导致整列数据被当成单个输入处理,最终编码出异常值。
  • 张量处理错误:squeeze(0)会错误压缩批量数据的维度,破坏输入结构。

修正后的代码示例

假设你的DataFrame有一列profile存储完整的句子文本:

from transformers import RobertaForSequenceClassification, RobertaTokenizer
import pandas as pd
import torch

model_name = "roberta-base"
model = RobertaForSequenceClassification.from_pretrained(model_name)
tokenizer = RobertaTokenizer.from_pretrained(model_name)

def encode_df(dataframe):
    # 提取所有profile文本,批量编码
    texts = dataframe["profile"].tolist()
    encoding = tokenizer.batch_encode_plus(
        texts,
        add_special_tokens=True,
        padding="longest",
        truncation=True,
        return_tensors="pt"
    )
    return encoding["input_ids"]

# 读取数据
df = pd.read_excel("/content/profile sentences into df.xlsx")
profiles_input = encode_df(df)

# 导出ONNX
torch.onnx.export(
    model,
    profiles_input,
    ONNX_path,
    input_names=['input'],
    output_names=['output'],
    dynamic_axes={
        'input': {0: 'batch_size', 1: 'sentence_length'},
        'output': {0: 'batch_size'}
    }
)

额外说明

  • 之前多列DataFrame的“正常”是巧合:零散字段被错误编码但未出现明显异常值,实际编码逻辑同样不正确。
  • 若你是拆分句子为多列存储(对应原fullName、Location等字段),需先拼接成完整句子再编码:
def combine_profile(row):
    return f"My name is {row['fullName']}. I live in {row['Location']}. I went to {row['College']} in {row['Degree_location']}. I have a degree in {row['Degree']}. I have worked as {row['Jobs']} in companies {row['Company']}. I have worked from {row['Date']}. My job locations have been {row['LocationJob']}. The average distance between my locations is {row['Avg_rounded']}."

df["profile"] = df.apply(combine_profile, axis=1)

内容的提问来源于stack exchange,提问作者Jesper Ezra

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.05 03:38:29