如何在Python中计算忽略位置的单词序列匹配百分比
单词序列匹配百分比计算方案
给定以下句子序列:
sentence1 = 'Ram is eating' sentence2 = 'is Ram eating' sentence3 = 'is Ram playing' sentence4 = 'movie Ram watching is'
由于difflib.SequenceMatcher是按字符匹配,无法直接满足忽略单词位置的匹配需求,我们可以通过以下方法实现指定规则的匹配百分比计算:
计算规则
匹配百分比公式:
match% = (sentence1与sentence2中匹配的单词数量/sentence1的总单词数)*100
实现步骤
- 将句子按空白字符拆分单词(自动处理连续空格)
- 通过集合求交集,得到两个句子共有的单词数量
- 代入公式计算百分比
代码实现
sentence1 = 'Ram is eating' sentence2 = 'is Ram eating' sentence3 = 'is Ram playing' sentence4 = 'movie Ram watching is' def get_match_percent(s1, s2): words1 = s1.split() words2 = s2.split() # 计算共同单词数量 common_words = set(words1) & set(words2) matched_count = len(common_words) total_words = len(words1) if total_words == 0: return 0.0 return (matched_count / total_words) * 100 # 输出结果 total_s1 = len(sentence1.split()) match1_2 = len(set(sentence1.split()) & set(sentence2.split())) print(f"sentence1与sentence2的匹配百分比 = {match1_2}/{total_s1} 即{get_match_percent(sentence1, sentence2):.2f}%") match1_3 = len(set(sentence1.split()) & set(sentence3.split())) print(f"sentence1与sentence3的匹配百分比 = {match1_3}/{total_s1} 即{get_match_percent(sentence1, sentence3):.2f}%") match1_4 = len(set(sentence1.split()) & set(sentence4.split())) print(f"sentence1与sentence4的匹配百分比 = {match1_4}/{total_s1} 即{get_match_percent(sentence1, sentence4):.2f}%")
运行结果
sentence1与sentence2的匹配百分比 = 3/3 即100.00% sentence1与sentence3的匹配百分比 = 2/3 即66.67% sentence1与sentence4的匹配百分比 = 2/3 即66.67%
内容的提问来源于stack exchange,提问作者k_p
相关产品推荐
相关产品推荐

