You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Omnivore模型动作识别:提取动词/名词分数结果异常求助

问题

我正在使用Omnivore模型,模型输出为3806个动宾结构动作的分数数组。我需要提取动词分数数组和名词分数数组,参考论文中“通过对动词边缘化得到名词预测,反之亦然”的描述,采用了边际概率方法,代码如下:

verb_scores_dict = {}
noun_scores_dict = {}

for pred, score in pred_class_names_all:
    verb, noun = pred.split()

if verb in verb_scores_dict:
  verb_scores_dict[verb] += score
else:
  verb_scores_dict[verb] = score

if noun in noun_scores_dict:
  noun_scores_dict[noun] += score
else:
  noun_scores_dict[noun] = score
      
result_dict={}
result_dict['verb_output'] = []
result_dict['noun_output'] = []
result_dict['narration_id'] = narration_id

# Print the verb and noun scores dictionaries
for verb, scores in verb_scores_dict.items():
    print(f"Verb: {verb}, Verb Scores: {scores}")
    result_dict['verb_output'].append(scores.tolist())
    

for noun, scores in noun_scores_dict.items():
    print(f"Noun: {noun}, Noun Scores: {scores}")
    result_dict['noun_output'].append(scores.tolist())

预期结果应为类似如下的数组:

{'verb_output': array([[ 7.72634077,  7.24404097,  0.4986968 , ..., -4.73686218,
        -4.68037271, -4.78593493],
       [10.65037537,  6.22874689,  2.58304691, ..., -4.34611273,
        -4.10662127, -3.87795758],
       [ 4.48676252,  8.9562788 ,  4.78455114, ..., -4.37370729,
        -5.02590704, -3.35811305],
       ...,
       [ 6.71048832,  8.05608559,  5.31442881, ..., -4.34917974,
        -5.15016031, -4.63199329],
       [ 7.93692255,  6.86483002,  3.15329337, ..., -3.71181631,
        -4.52799511, -4.36749268],
       [ 8.39965153,  9.5401535 , 15.13584709, ..., -6.13945246,
        -5.81251431, -3.81169009]]), 'noun_output': array([[-6.6729846 , -2.1299243 ,  1.14122081, ..., -4.76375723,
        -3.65956378, -4.85687351],
       [-0.07463244,  3.85092378,  6.30206203, ..., -4.32961893,
        -3.71707988, -3.77974272],
       [ 3.05986619,  1.97096956,  2.50316858, ..., -4.64014769,
        -3.84915638, -3.59207392],
       ...,
       [ 4.28227949,  3.72166443,  7.30290747, ..., -5.12706661,
        -4.40550518, -3.26095915],
       [ 3.11553764,  7.43528414,  5.56893921, ..., -4.48091173,
        -4.32157898, -2.97596955],
       [ 6.16293764,  3.73808503,  6.44398689, ..., -6.22030973,
        -4.37711859, -5.48904181]]), 'narration_id': array(['P01_101_0', 'P01_101_1', 'P01_101_10', ..., 'P33_105_658',
       'P33_105_659', 'P33_105_66'], dtype='<U11')}

但实际得到的结果数值过大且多为负数:

{'verb_output': array([[ -43.78885651, -139.56182861, -110.56599426, ...,  -12.32396984,
          -8.73503685,  -15.43686867],
       [ -42.33675766, -108.72449493, -187.36923218, ...,  -12.29256916,
          -9.07168865,  -15.52522469],
       [ -51.75484848, -119.52777863,  -63.90496063, ...,  -10.2333765 ,
         -14.77134037,   -2.52343321],
       ...,
       [-493.74246216, -503.70645142, -207.50227356, ...,   -5.77831984,
         -22.88568878,   -7.9034214 ],
       [-237.2469635 , -579.89691162, -491.43731689, ...,  -16.16319847,
          -7.59590626,   -8.50958157],
       [-172.81614685, -318.99887085, -426.69381714, ...,   -5.8961854 ,
         -16.30410767,   -8.4819355 ]]), 'noun_output': array([[ -43.78885651, -139.56182861,  -14.37578964, ...,   -2.71890616,
          -2.72560787,  -15.99341011],
       [ -42.33675766, -187.36923218,  -33.00939178, ...,   -5.39493179,
         -15.75678635,   -2.62935948],
       [ -51.75484848, -266.3885498 , -440.42184448, ...,  -15.25948334,
          -5.51948929,   -3.1326921 ],
       ...,
       [-493.74246216, -503.70645142, -207.50227356, ...,   -7.68950129,
          -3.64070535,   -7.95105267],
       [-237.2469635 , -579.89691162, -491.43731689, ...,   -4.07324791,
          -8.35855579,   -4.86216784],
       [-172.81614685, -318.99887085, -426.69381714, ...,   -7.36990547,
          -7.62912035,   -3.9157083 ]]), 'narration_id': array(['P01_11_100', 'P01_11_101', 'P01_11_102', ..., 'P28_26_0',
       'P28_26_10', 'P28_26_1'], dtype='<U10')}

请问是否需要归一化处理,或是该方法不适用于分数提取?


分析与解决方案

  1. 修复代码缩进错误
    你的循环内的动词、名词处理逻辑没有被包含在for循环块中,导致只会处理最后一组pred, score,而非遍历所有3806个动宾对。修正后的循环代码如下:

    for pred, score in pred_class_names_all:
        verb, noun = pred.split()
        if verb in verb_scores_dict:
            verb_scores_dict[verb] += score
        else:
            verb_scores_dict[verb] = score
        if noun in noun_scores_dict:
            noun_scores_dict[noun] += score
        else:
            noun_scores_dict[noun] = score
    
  2. 理解模型输出的分数性质
    Omnivore输出的分数是logits(未经过softmax的原始分数),不是概率值。直接累加logits会导致数值异常,因为logits可以是任意正负值,累加后会放大偏差。论文中提到的“边缘化”操作,是基于概率分布的,而非原始logits。

  3. 正确的边缘化处理流程

    • 先对每个动宾对的logits做softmax转换为概率值,确保所有动宾对的概率和为1;
    • 对同一个动词的所有动宾对概率求和,得到该动词的边际概率,同理处理名词;
    • 如果需要保持类似预期结果的分数形式,可以对边际概率取log转换回对数空间,但要注意加极小值epsilon避免log(0)的数值错误。
  4. 归一化的必要性
    归一化是必要的,但不是直接对累加后的结果归一化,而是先将logits转换为概率后再进行边缘化计算。如果直接使用logits做边缘化,也可以对每个样本的logits先做log-sum-exp处理,再进行累加,避免数值溢出。

内容的提问来源于stack exchange,提问作者Georgia

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.13 08:17:33