You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Java句子相似度计算:求相同词数占长句词数的比例及所属类型

问题解答

一、相似度类别

你提到的这种相似度计算方式属于基于单词集合的重叠相似度,是Jaccard相似度的简化变体。标准Jaccard相似度以两个单词集合的交集大小除以并集大小,而你这里用交集大小除以较长句子的单词总数,核心是衡量较短句子的单词在较长句子中的覆盖比例。

二、实现方案(无需额外库,手动实现更高效)

这种逻辑非常简单,不需要依赖第三方库,直接手动实现即可。以下是完整Java代码:

import java.util.Arrays;
import java.util.HashSet;
import java.util.Set;

public class SentenceSimilarity {
    public static void main(String[] args) {
        String sentence1 = "Jack go to basketball.";
        String sentence2 = "Jack go to basketball match";
        
        // 预处理:清理标点、转小写(可根据需求去掉toLowerCase()区分大小写)、转为单词集合
        Set<String> words1 = getWordSet(sentence1);
        Set<String> words2 = getWordSet(sentence2);
        
        // 计算相同单词数量(交集大小)
        Set<String> intersection = new HashSet<>(words1);
        intersection.retainAll(words2);
        int commonWordsCount = intersection.size();
        
        // 获取较长句子的单词总数
        int maxWordCount = Math.max(words1.size(), words2.size());
        
        // 计算相似度比例,避免除以0的情况
        double similarity = maxWordCount == 0 ? 0.0 : (double) commonWordsCount / maxWordCount;
        
        System.out.printf("相似度比例:%.2f%n", similarity); // 示例输出:相似度比例:0.80
    }
    
    private static Set<String> getWordSet(String sentence) {
        // 移除非字母和空格的字符,按空格分割单词
        String cleanedSentence = sentence.replaceAll("[^a-zA-Z\\s]", "").toLowerCase();
        return new HashSet<>(Arrays.asList(cleanedSentence.split("\\s+")));
    }
}

三、可选第三方库

如果想用现有库简化集合操作,可以选择:

  • Google Guava:提供Sets.intersection()方法直接获取交集
  • Apache Commons Collections:提供CollectionUtils.intersection()方法计算交集

以Guava为例的简化实现:

import com.google.common.collect.Sets;
import java.util.Arrays;
import java.util.Set;

public class SentenceSimilarityWithGuava {
    public static void main(String[] args) {
        String sentence1 = "Jack go to basketball.";
        String sentence2 = "Jack go to basketball match";
        
        Set<String> words1 = getWordSet(sentence1);
        Set<String> words2 = getWordSet(sentence2);
        
        int commonWordsCount = Sets.intersection(words1, words2).size();
        int maxWordCount = Math.max(words1.size(), words2.size());
        double similarity = maxWordCount == 0 ? 0.0 : (double) commonWordsCount / maxWordCount;
        
        System.out.printf("相似度比例:%.2f%n", similarity);
    }
    
    private static Set<String> getWordSet(String sentence) {
        String cleanedSentence = sentence.replaceAll("[^a-zA-Z\\s]", "").toLowerCase();
        return Sets.newHashSet(Arrays.asList(cleanedSentence.split("\\s+")));
    }
}

使用Guava需引入Maven依赖(版本可替换为最新稳定版):

<dependency>
    <groupId>com.google.guava</groupId>
    <artifactId>guava</artifactId>
    <version>32.1.3-jre</version>
</dependency>

内容的提问来源于stack exchange,提问作者Emre Terzi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 02:41:03