You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

多语言影片标题匹配问题:含&及多语言‘和’等价词的DataFrame合并

解决多语言影片标题匹配问题的实用方案

针对不同语言影片标题无法直接匹配、又不能依赖停用词或硬编码替换的问题,这里提供几个无需提前知晓语言列表、不会过度修改原标题的可行方案:

方案1:基于词袋集合的Jaccard相似度匹配

核心思路是聚焦标题里的核心名词,弱化连接词(&、e、und这类)的干扰。通过将标题拆解为单词集合,计算两个集合的Jaccard相似度(交集大小/并集大小),超过设定阈值就判定为匹配项。

实现步骤与代码:

import pandas as pd
from itertools import product
import string

# 示例数据
df1 = pd.DataFrame({'film': ['Beavis & Butthead', 'Bonnie e Clyde', 'Adam & Eve']})
df2 = pd.DataFrame({'film': ['Beavis und Butthead', 'Bonnie & Clyde', 'Adam et Eve']})

# 标题预处理:去标点、转小写、过滤短词(减少连接词干扰)
def preprocess_title(title):
    title_clean = title.translate(str.maketrans('', '', string.punctuation))
    words = title_clean.lower().split()
    return set([word for word in words if len(word) > 2])

# 对两个DataFrame的标题做预处理
df1['processed'] = df1['film'].apply(preprocess_title)
df2['processed'] = df2['film'].apply(preprocess_title)

# 生成所有标题配对并计算相似度
matches = []
for row1, row2 in product(df1.itertuples(), df2.itertuples()):
    intersection = len(row1.processed & row2.processed)
    union = len(row1.processed | row2.processed)
    if union == 0:
        continue
    jaccard = intersection / union
    if jaccard >= 0.7:  # 可根据实际数据调整阈值
        matches.append({
            'film_df1': row1.film,
            'film_df2': row2.film,
            'similarity': jaccard
        })

# 合并原DataFrame
match_df = pd.DataFrame(matches)
final_merge = pd.merge(df1, match_df, left_on='film', right_on='film_df1').merge(df2, left_on='film_df2', right_on='film')

方案2:使用字符串模糊匹配库(RapidFuzz)

这类库专门处理字符串相似度,其中token_set_ratio方法会将字符串拆分为单词集合,自动忽略顺序和重复内容,完美适配核心单词一致但连接词不同的场景,且无需依赖语言信息。

代码示例:

import pandas as pd
from rapidfuzz import process, fuzz

# 示例数据
df1 = pd.DataFrame({'film': ['Beavis & Butthead', 'Bonnie e Clyde', 'Adam & Eve']})
df2 = pd.DataFrame({'film': ['Beavis und Butthead', 'Bonnie & Clyde', 'Adam et Eve']})

# 为df1的每个标题在df2中匹配最相似的结果
matches = []
for title in df1['film']:
    # 用token_set_ratio计算相似度,设置匹配阈值为80
    result = process.extractOne(title, df2['film'], scorer=fuzz.token_set_ratio, score_cutoff=80)
    if result:
        matches.append({
            'film_df1': title,
            'film_df2': result[0],
            'similarity_score': result[1]
        })

# 合并原数据
match_df = pd.DataFrame(matches)
final_merge = pd.merge(df1, match_df, left_on='film', right_on='film_df1').merge(df2, left_on='film_df2', right_on='film')

方案3:基于多语言文本嵌入模型

如果标题结构复杂(比如包含副标题、跨语言核心词变体),可以用预训练的多语言文本嵌入模型,将标题转为语义向量后计算余弦相似度,能捕捉深层语义关联,完全不需要提前知道语言类型。

代码示例:

import pandas as pd
from sentence_transformers import SentenceTransformer, util

# 加载支持100+语言的预训练模型
model = SentenceTransformer('all-MiniLM-L6-v2')

# 示例数据
df1 = pd.DataFrame({'film': ['Beavis & Butthead', 'Bonnie e Clyde', 'Adam & Eve']})
df2 = pd.DataFrame({'film': ['Beavis und Butthead', 'Bonnie & Clyde', 'Adam et Eve']})

# 生成标题的语义嵌入向量
embeddings1 = model.encode(df1['film'].tolist(), convert_to_tensor=True)
embeddings2 = model.encode(df2['film'].tolist(), convert_to_tensor=True)

# 计算余弦相似度并筛选匹配项
cosine_scores = util.cos_sim(embeddings1, embeddings2)
matches = []
for i in range(len(df1)):
    for j in range(len(df2)):
        if cosine_scores[i][j] >= 0.8:  # 阈值可根据需求调整
            matches.append({
                'film_df1': df1.iloc[i]['film'],
                'film_df2': df2.iloc[j]['film'],
                'cosine_similarity': float(cosine_scores[i][j])
            })

# 合并原数据
match_df = pd.DataFrame(matches)
final_merge = pd.merge(df1, match_df, left_on='film', right_on='film_df1').merge(df2, left_on='film_df2', right_on='film')

方案选择建议:

  • 数据量小、标题结构简单:优先选方案1或2,速度快且易调整
  • 数据量大、标题复杂或涉及多语言变体:选方案3,语义匹配准确性更高

内容的提问来源于stack exchange,提问作者Azamat Bagatov

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.29 01:07:50