You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求可检测两个文件内容相似度百分比的Java插件,附现有Maven依赖配置

Java Tools for Calculating File Content Similarity Percentage

Hey there! Since you're already working with string similarity libraries like net.ricecode:string-similarity and info.debatty:java-string-similarity, let's cover your options for checking file content similarity in Java—either by extending what you already have, or using specialized tools.

Option 1: Extend Your Existing String Similarity Libraries

You don't need a brand-new plugin to get file similarity checks. You can easily adapt the libraries you're already using by first reading file contents into strings. This works great for most text-based files (TXT, JSON, XML, etc.).

Here's a quick example using the Levenshtein Distance from your current dependencies:

import java.nio.file.Files;
import java.nio.file.Paths;
import net.ricecode.string.similarity.LevenshteinDistance;

public class FileSimilarityCalculator {
    public static void main(String[] args) throws Exception {
        // Read file contents into strings
        String content1 = Files.readString(Paths.get("/path/to/your/first/file"));
        String content2 = Files.readString(Paths.get("/path/to/your/second/file"));

        // Calculate similarity using Levenshtein Distance
        LevenshteinDistance distanceCalculator = new LevenshteinDistance();
        int editDistance = distanceCalculator.distance(content1, content2);
        
        // Convert edit distance to a percentage similarity
        int maxLength = Math.max(content1.length(), content2.length());
        double similarityPercent = maxLength == 0 ? 100.0 : (1.0 - (double) editDistance / maxLength) * 100;

        System.out.printf("Files are %.2f%% similar%n", similarityPercent);
    }
}

For larger files, use the SimHash algorithm from info.debatty:java-string-similarity—it generates hash values for file content (you can process chunks to avoid loading the entire file into memory) and calculates similarity via Hamming distance, which is far more memory-efficient.

Option 2: Specialized Libraries for File Similarity

If you want tools tailored specifically for file-level comparison (especially for code files), these are worth checking out:

PMD Copy-Paste Detector (CPD)

Perfect for source code (Java, Python, etc.)—it identifies duplicated code blocks and metrics based on code structure, not just raw text. It ignores whitespace and comments, making it more accurate for code than plain string comparisons.

Add these Maven dependencies:

<dependency>
    <groupId>net.sourceforge.pmd</groupId>
    <artifactId>pmd-core</artifactId>
    <version>7.0.0</version>
</dependency>
<dependency>
    <groupId>net.sourceforge.pmd</groupId>
    <artifactId>pmd-java</artifactId>
    <version>7.0.0</version>
</dependency>

A basic usage example to check two Java files:

import net.sourceforge.pmd.cpd.CPD;
import net.sourceforge.pmd.cpd.CPDConfiguration;
import net.sourceforge.pmd.cpd.JavaLanguage;
import java.io.File;
import java.util.List;

public class CodeSimilarityChecker {
    public static void main(String[] args) {
        CPDConfiguration config = new CPDConfiguration();
        config.setLanguage(new JavaLanguage());
        config.setMinimumTileSize(10); // Minimum tokens to consider a duplicate block

        CPD cpd = new CPD(config);
        cpd.add(new File("/path/to/FirstClass.java"));
        cpd.add(new File("/path/to/SecondClass.java"));

        cpd.go();
        List<?> duplicates = cpd.getDuplicates();
        
        if (duplicates.isEmpty()) {
            System.out.println("No duplicated code found.");
        } else {
            duplicates.forEach(dup -> {
                System.out.println("Found duplicate block:");
                System.out.println("Start line in file 1: " + dup.getMarkers().get(0).getBeginLine());
                System.out.println("Start line in file 2: " + dup.getMarkers().get(1).getBeginLine());
                System.out.println("Duplicate lines count: " + dup.getLineCount());
                // Calculate percentage similarity using total file lines vs duplicated lines
            });
        }
    }
}

Apache Commons IO (For Exact Matches + Diffing)

While it doesn't output a percentage directly, Apache Commons IO has FileUtils.contentEquals() to check if two files are identical. For granular diffs, combine it with Apache Commons Text's StringDiffBuilder to generate a diff report, then calculate similarity manually from the results.

Final Notes

  • For plain text files: Stick with your existing string similarity libraries + file reading—simple and effective.
  • For large files: Use SimHash from info.debatty:java-string-similarity to avoid memory overload.
  • For source code: PMD CPD is the most robust option, as it understands code syntax and avoids false positives from irrelevant whitespace/comments.

内容的提问来源于stack exchange,提问作者dijo francis

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 08:05:08