You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Java向HashMap<String,String>添加唯一值:CSV仅写入新文件异常排查

Hey there! Let's break down why your MD5 matching isn't working and fix that "only write new files" logic.

1. First, Diagnose the CSV Format & Parsing Issue

Looking at your CSV snippet, the structure seems mixed up: 4d1954a6d4e99cacc57beef94c80f994,uiautomationcoreapi.h;E:\Tools\Strawberry-perl-5.24.1.1-64\c\x... uses both commas and semicolons as separators. This is likely throwing off your ability to correctly extract and match MD5 values later.

CSV files rely on consistent delimiters—stick to a single separator (like commas) for each row, e.g.:

4d1954a6d4e99cacc57beef94c80f994,E:\Tools\Strawberry-perl-5.24.1.1-64\c\x\...\uiautomationcoreapi.h
2. Fix the MD5 Matching Core Logic

The key issue is almost certainly how you're storing and checking existing MD5 hashes. Here's a step-by-step fix:

  • Load existing MD5s into a Set: When your program starts, read the CSV and store all existing MD5 values in a HashSet<String>—this makes lookups fast and reliable.
  • Normalize case for consistency: MD5 hashes are often written in lowercase, but if your code generates uppercase hashes (or vice versa), the match will fail. Force all hashes to the same case (lowercase is standard) when storing and checking.
  • Trim whitespace: Accidental newlines or spaces in your CSV rows can make MD5 values appear mismatched even if they're identical. Use trim() to clean up each hash when reading.
3. Sample Code Snippets to Implement the Fix

Loading Existing MD5 Hashes

Set<String> existingMd5Hashes = new HashSet<>();
File csvFile = new File("your-output.csv");

if (csvFile.exists()) {
    try (BufferedReader br = new BufferedReader(new FileReader(csvFile))) {
        String line;
        while ((line = br.readLine()) != null) {
            // Split each row into max 2 parts (MD5 and file path)
            String[] rowParts = line.split(",", 2);
            if (rowParts.length >= 1) {
                // Normalize to lowercase and trim whitespace
                String cleanMd5 = rowParts[0].trim().toLowerCase();
                existingMd5Hashes.add(cleanMd5);
            }
        }
    } catch (IOException e) {
        e.printStackTrace();
    }
}

Checking New Files & Writing to CSV

// For each file you detect in the directory:
String fileMd5 = calculateFileMd5(targetFile); // Your existing MD5 calculation method
String normalizedMd5 = fileMd5.toLowerCase();

if (!existingMd5Hashes.contains(normalizedMd5)) {
    // Write the new entry to CSV using the standard format
    writeToCsv(normalizedMd5, targetFile.getAbsolutePath());
    // Add the hash to the set to avoid duplicate writes in the same run
    existingMd5Hashes.add(normalizedMd5);
}
4. Common Pitfalls to Double-Check
  • File content vs. metadata: Ensure your calculateFileMd5 method computes the hash from the file's content, not its filename or modification time—using metadata would break matching for unchanged files.
  • CSV write consistency: When appending new rows to the CSV, make sure you're using the same delimiter (comma) and case for MD5 values as you do when reading.
  • Edge cases: Handle empty CSV files, files that can't be read (skip them instead of crashing), and duplicate files in different directories (your current logic uses MD5, so this is okay—just ensure the full path is stored).

内容的提问来源于stack exchange,提问作者DidierDR

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 09:50:43