You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中DataFrame逐行计算相似度存入新列的异常解决

问题分析

你遇到的问题核心是赋值逻辑错误:在循环里执行merged['fsim'] = fsim时,你是把整个fsim列的所有行都覆盖成当前循环迭代的相似度值,等循环跑完最后一次,所有行自然都会变成最后一次计算的结果,这就是为啥所有得分完全一致。另外注意你代码里的变量名小疏漏:计算得到的是fsims,但赋值用的是fsim,这会触发NameError,应该是笔误。

解决方案

下面提供两种可行的修正方式,优先推荐第二种更高效的方法:

方法1:修正循环内的赋值逻辑

在iterrows循环中,通过loc索引精准定位当前行的fsim列赋值,而不是覆盖整列:

# 先初始化fsim列,避免赋值时出错
merged['fsim'] = 0.0

for i, row in merged.iterrows():
    captions = [row['cap1'], row['cap2']]
    # 预处理两个caption
    for idx in range(len(captions)):
        captions[idx] = pre_process(captions[idx])
        captions[idx] = lemmatize_sentence(captions[idx])
    # 生成特征向量并计算相似度
    feature_vectors = tfidf_vectorizer.transform(captions)
    fsim_score = get_cosine_similarity(feature_vectors[0], feature_vectors[1])
    # 给当前行赋值
    merged.loc[i, 'fsim'] = fsim_score

方法2:使用apply函数(更高效,推荐)

Pandas的apply函数可以直接对每行数据批量处理,避免手动循环,代码更简洁且性能更优(尤其适合大数据集):

def calculate_fsim(row):
    # 依次处理两个caption
    cap1_processed = lemmatize_sentence(pre_process(row['cap1']))
    cap2_processed = lemmatize_sentence(pre_process(row['cap2']))
    # 生成特征向量并计算余弦相似度
    feature_vectors = tfidf_vectorizer.transform([cap1_processed, cap2_processed])
    return get_cosine_similarity(feature_vectors[0], feature_vectors[1])

# 直接生成fsim列
merged['fsim'] = merged.apply(calculate_fsim, axis=1)

额外提示

  • 确保tfidf_vectorizer已经用所有相关的caption数据完成拟合(比如tfidf_vectorizer.fit(all_captions)),否则生成的特征向量会失去参考意义。
  • 尽量避免用iterrows处理大规模数据,它的运行效率远低于向量化操作或apply。

内容的提问来源于stack exchange,提问作者Vaidehi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 16:57:43