Java读取CSV文件并检测重复项的最佳实现方案咨询
Hey there! I see you're stuck with inconsistent array lengths when parsing your CSV, which is making it impossible to reliably track duplicates. Let's break down why this is happening and fix it step by step.
Why Your Current Code Causes Inconsistent Array Lengths
The core issue here is using String.split(",") to parse CSV lines. CSV files often include fields with commas wrapped in quotes (like "New York, NY"), and your split method doesn't recognize that these commas are part of the field value—not separators. This splits valid single fields into multiple elements, leading to arrays of varying lengths across lines.
Fixes & Best Practices
1. Use a Reliable CSV Parsing Library (Recommended)
Instead of building your own CSV parser from scratch, use a battle-tested library like OpenCSV that handles edge cases (quoted fields, escaped characters, line breaks inside fields) automatically.
First, add the OpenCSV dependency (if using Maven):
<dependency> <groupId>com.opencsv</groupId> <artifactId>opencsv</artifactId> <version>5.6</version> </dependency>
Here's how to rewrite your code to read CSV correctly and track duplicates:
import com.opencsv.CSVReader; import java.io.FileReader; import java.io.IOException; import java.util.HashSet; import java.util.Set; public class CSVReaderWithDuplicateCheck { public static void main(String[] args) { String csvFile = "Cchallenge.csv"; Set<String> uniqueRecords = new HashSet<>(); Set<String> duplicateRecords = new HashSet<>(); try (CSVReader reader = new CSVReader(new FileReader(csvFile))) { String[] nextLine; while ((nextLine = reader.readNext()) != null) { // Convert the record array to a single string for easy comparison String recordKey = String.join("|", nextLine); // If add() returns false, the record already exists (duplicate) if (!uniqueRecords.add(recordKey)) { duplicateRecords.add(recordKey); } } } catch (IOException e) { e.printStackTrace(); } // Print out found duplicates System.out.println("Duplicate records detected:"); duplicateRecords.forEach(System.out::println); } }
2. Implement a Basic CSV Splitter (No Third-Party Libraries)
If you can't use external libraries, you can write a simple method to split CSV lines while ignoring commas inside quotes. Here's a lightweight implementation:
import java.io.BufferedReader; import java.io.FileReader; import java.io.IOException; import java.util.ArrayList; import java.util.HashSet; import java.util.List; import java.util.Set; public class CustomCSVReader { public static void main(String[] args) { String csvFile = "Cchallenge.csv"; Set<String> uniqueRecords = new HashSet<>(); Set<String> duplicateRecords = new HashSet<>(); try (BufferedReader br = new BufferedReader(new FileReader(csvFile))) { String line; while ((line = br.readLine()) != null) { List<String> fields = splitCSVLine(line); String recordKey = String.join("|", fields); if (!uniqueRecords.add(recordKey)) { duplicateRecords.add(recordKey); } } } catch (IOException e) { e.printStackTrace(); } System.out.println("Duplicate records detected:"); duplicateRecords.forEach(System.out::println); } private static List<String> splitCSVLine(String line) { List<String> fields = new ArrayList<>(); StringBuilder currentField = new StringBuilder(); boolean inQuotes = false; for (char c : line.toCharArray()) { if (c == '"') { // Toggle quote state inQuotes = !inQuotes; } else if (c == ',' && !inQuotes) { // Only split on commas outside quotes fields.add(currentField.toString().trim()); currentField.setLength(0); } else { currentField.append(c); } } // Add the final field after loop ends fields.add(currentField.toString().trim()); return fields; } }
Key Tips for Duplicate Tracking
- Use
HashSetfor unique records: It has O(1) average time complexity for add/contains operations, making duplicate checks efficient. - Customize the record key: If you only need to track duplicates of specific fields (not the entire record), create a key from those fields instead of the whole line.
- Choose a safe separator: When joining fields into a record key, use a character that doesn't appear in your CSV data (like
|or;) to avoid false duplicates.
内容的提问来源于stack exchange,提问作者John Gorgy

