Python遍历DataFrame实现系列电影剧情与原作/前作相似度比对
实现代码及步骤
前置依赖导入
import pandas as pd from sklearn.feature_extraction.text import CountVectorizer from scipy.spatial import distance
相似度计算函数优化
原函数将依赖导入放在函数内部会导致批量调用时重复导入、效率降低,优化后版本如下,同时新增空文本异常处理避免报错:
def cosine_distance_countvectorizer_method(s1, s2): # 空文本直接返回空值 if pd.isna(s1) or pd.isna(s2) or not s1.strip() or not s2.strip(): return None allsentences = [s1 , s2] vectorizer = CountVectorizer() all_sentences_to_vector = vectorizer.fit_transform(allsentences) v1 = all_sentences_to_vector.toarray()[0] v2 = all_sentences_to_vector.toarray()[1] cosine = distance.cosine(v1, v2) # 如需直接返回0-1范围的相似度可修改为 return 1 - cosine return cosine
核心比对逻辑实现
假设你的数据集存储在DataFrame df中,包含FranID(系列ID)、Seq(序列编号)、plot(剧情文本)三个关键字段,可直接运行下述代码生成结果列:
# 先按系列ID、序列编号排序,避免序列混乱导致比对错误 df = df.sort_values(by=['FranID', 'Seq']).reset_index(drop=True) # 逻辑1:同系列所有作品和首部(Seq=0)比对 def calc_first_sim(group): first_plot = group.loc[group['Seq'] == 0, 'plot'].iloc[0] if len(group.loc[group['Seq'] == 0]) > 0 else None group['sim_with_first'] = group['plot'].apply(lambda x: cosine_distance_countvectorizer_method(x, first_plot)) # 首部自身的比对结果设为空 group.loc[group['Seq'] == 0, 'sim_with_first'] = None return group df = df.groupby('FranID', group_keys=False).apply(calc_first_sim) # 逻辑2:同系列每部作品和上一部作品比对 def calc_prev_sim(group): group['prev_plot'] = group['plot'].shift(1) group['sim_with_previous'] = group.apply(lambda row: cosine_distance_countvectorizer_method(row['plot'], row['prev_plot']), axis=1) group = group.drop('prev_plot', axis=1) return group df = df.groupby('FranID', group_keys=False).apply(calc_prev_sim)
结果说明
sim_with_first列存储当前作品和同系列首部的余弦距离,如需转换为百分比相似度可计算(1 - df['sim_with_first'])*100sim_with_previous列存储当前作品和同系列前一部的余弦距离,Seq为0的行该字段为空- 如你的剧情文本列名不是
plot,替换代码中对应位置的字段名即可
内容的提问来源于stack exchange,提问作者bzh
相关产品推荐
相关产品推荐

