You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从字典值中获取键并映射到DataFrame列中

问题描述

现有一个名为df_x_encode的DataFrame,结构如下:

x_test_encode
0   [0.1260023, -0.014597204, -0.079445906, -0.055...
1   [0.0083509395, 0.09799187, -0.05743032, -0.000...
2   [-0.05807189, 0.11802298, -0.031580053, -0.064...
3   [0.1260023, -0.014597204, -0.079445906, -0.055...
4   [0.121216424, -0.017603464, -0.090226464, -0.0...

同时有一个字典(假设名为vec_to_text),键为文本内容,值是与x_test_encode列匹配的向量数组,示例如下:

{'Strengthening the field is a must ': np.array([ 1.75993890e-02,  7.26785734e-02, -7.36519024e-02, -2.17226259e-02,
         3.65523808e-02, -4.50823084e-03,  6.18522726e-02,  1.35725755e-02,
        -1.65322982e-02, -1.93105303e-02, -6.45413473e-02, -1.43367276e-02,
         3.43437083e-02, -5.04908897e-02, -7.43871846e-04, -2.44313944e-02,
         2.88490783e-02, -2.72445306e-02,  5.23326918e-02,  4.61216345e-02,
         2.41497066e-04, -8.29233676e-02, -9.53390170e-03, -7.67266843e-03,..]),
...

需要为DataFrame新增text列,根据x_test_encode中的向量匹配字典里的对应文本,期望效果如下:

x_test_encode                                        text
0   [0.1260023, -0.014597204, -0.079445906, -0.055...    This is to be noted that..
1   [0.0083509395, 0.09799187, -0.05743032, -0.000...    Strengthening the perfect..
2   [-0.05807189, 0.11802298, -0.031580053, -0.064...   
3   [0.1260023, -0.014597204, -0.079445906, -0.055...
4   [0.121216424, -0.017603464, -0.090226464, -0.0...
解决方案

由于列表或numpy数组是不可哈希类型,无法直接作为字典的键,需要先反转字典并将向量转为可哈希的元组,再进行映射:

步骤1:反转字典,将向量转为元组作为键

import numpy as np

# 反转原字典,把向量转成元组当键,文本作为值
reverse_dict = {tuple(vec): text for text, vec in vec_to_text.items()}
  • 如果原字典的值是numpy数组,直接转元组即可;如果是列表,同样用tuple(vec)转换。

步骤2:为DataFrame新增匹配列

# 对x_test_encode列的每个向量转元组,从反转字典中匹配对应文本
df_x_encode['text'] = df_x_encode['x_test_encode'].apply(
    lambda x: reverse_dict.get(tuple(np.array(x)), '')
)
  • 若x_test_encode列的元素本身就是numpy数组,可简化为tuple(x);若为列表则直接tuple(x)。

处理精度匹配问题

如果存在浮点精度差异导致匹配失败,可以对向量做四舍五入后再转元组:

# 保留6位小数后构建反转字典
reverse_dict = {tuple(np.round(vec, 6)): text for text, vec in vec_to_text.items()}

# 同样对DataFrame中的向量做四舍五入后匹配
df_x_encode['text'] = df_x_encode['x_test_encode'].apply(
    lambda x: reverse_dict.get(tuple(np.round(np.array(x), 6)), '')
)

内容的提问来源于stack exchange,提问作者d_b

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 18:35:15