You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将自有数据集与Hugging Face的wiki_dpr合并并匹配其Embedding方式?

生成与wiki_dpr数据集一致的Embedding解决方案
  • 核心问题:你用错了模型。wiki_dpr自带的Embedding是通过facebook/dpr-ctx_encoder-multiset-base生成的,而非你当前使用的facebook/dpr-ctx_encoder-single-nq-base,这是导致结果差异的关键原因。

  • 修正后的代码:

!pip install datasets evaluate transformers[sentencepiece]
!apt install libomp-dev
!pip install faiss-cpu
!pip install -U sentence-transformers
from transformers import DPRContextEncoder, DPRContextEncoderTokenizer

# 数据
add_dpr_index = "7"
add_dpr_title = "Aaron"
add_dpr_text = "in literature dating to the Babylonian captivity and later. The books of Judges, Samuel and Kings mention priests and Levites, but do not mention the Aaronides in particular. The Book of Ezekiel, which devotes much attention to priestly matters, calls the priestly upper class the Zadokites after one of King David's priests. It does reflect a two-tier priesthood with the Levites in subordinate position. A two-tier hierarchy of Aaronides and Levites appears in Ezra, Nehemiah and Chronicles. As a result, many historians think that Aaronide families did not control the priesthood in pre-exilic Israel. What is clear is that high"

# 替换为wiki_dpr对应的模型
tokenizer = DPRContextEncoderTokenizer.from_pretrained("facebook/dpr-ctx_encoder-multiset-base")
model = DPRContextEncoder.from_pretrained("facebook/dpr-ctx_encoder-multiset-base")
input_ids = tokenizer(add_dpr_text, return_tensors="pt")["input_ids"]  # 修正变量名拼写错误
embeddings = model(input_ids).pooler_output
print(embeddings)
  • 额外提示:你原代码里存在变量名拼写错误,add_dprtext应改为add_dpr_text,修正后才能正确读取目标文本。

内容的提问来源于stack exchange,提问作者강동근

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.23 02:37:50