You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何对pandas DataFrame所有值应用Bert编码获取嵌入向量

报错原因

applymap的参数需要是接收单个单元格值作为入参的可调用函数,原代码里直接写model.encode(convert_to_tensor=True)是立刻执行了encode方法,没有传入方法要求的必填位置参数sentences,所以触发参数缺失报错。

实现方案

根据数据量大小可以选两种实现方式:

  • 小数据集场景:用lambda匿名函数包装encode调用,逐元素传入文本编码
  • 大数据集场景:先展平所有文本批量编码,再重构为原DataFrame结构,推理速度远高于逐元素调用

逐元素编码实现(小数据量用)

from sentence_transformers import SentenceTransformer
import pandas as pd

model = SentenceTransformer('sentence-transformers/all-MiniLM-L6-v2')
sentences = [["I'm happy", "I'm full of happiness"], ["I am sam", "I am good"]]
df = pd.DataFrame(sentences)

# 给applymap传入可调用的lambda,每个单元格的值x作为待编码文本传入
encoded_df = df.applymap(lambda x: model.encode(x, convert_to_tensor=True))

批量编码实现(推荐,性能更好)

逐单元格重复调用encode无法触发模型的批量推理优化,文本量较大时速度很慢,推荐用批量编码的方式:

from sentence_transformers import SentenceTransformer
import pandas as pd

model = SentenceTransformer('sentence-transformers/all-MiniLM-L6-v2')
sentences = [["I'm happy", "I'm full of happiness"], ["I am sam", "I am good"]]
df = pd.DataFrame(sentences)

# 展平所有文本为一维列表
all_text = df.values.flatten().tolist()
# 一次性完成所有文本的编码
all_embeddings = model.encode(all_text, convert_to_tensor=True)
# 将编码结果还原为原DataFrame的行列结构
encoded_df = pd.DataFrame(
    all_embeddings.reshape(df.shape[0], df.shape[1], -1).tolist(),
    index=df.index,
    columns=df.columns
)

两种方式最终得到的encoded_df中,每个单元格存储的就是对应位置文本的BERT嵌入向量。

内容的提问来源于stack exchange,提问作者usr_lal123

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.28 05:51:17