Java读取格式错误的CSV数据时出现异常,求正确处理方案
Hey, that's a super common pain point with messy CSV files—missing closing quotes can totally break standard parsers, which is exactly what's happening here with that first record throwing the Invalid char between encapsulated token and delimiter error. Let's walk through practical solutions to read this kind of CSV in Java, using both popular libraries and a manual approach if you want to avoid adding dependencies.
Option 1: Apache Commons CSV (Recommended for Flexibility)
Apache Commons CSV is my go-to for handling edge-case CSVs because it has tons of configurable options. Here's how to use it to work around unclosed quotes:
First, add the Maven dependency (adjust the version to the latest if needed):
<dependency> <groupId>org.apache.commons</groupId> <artifactId>commons-csv</artifactId> <version>1.10.0</version> </dependency>
Then, the code:
import org.apache.commons.csv.CSVFormat; import org.apache.commons.csv.CSVParser; import org.apache.commons.csv.CSVRecord; import org.apache.commons.csv.QuoteMode; import java.io.StringReader; import java.util.List; public class BrokenCsvHandler { public static void main(String[] args) { String wonkyCsv = "1,\"Robert,Adams,Washington,US\n2,\"Madhu\",Grant,Oregon,US\n3,Mohan,Young,Texas,US"; // Set up a lenient CSV format that tolerates missing column names and minimal quoting CSVFormat csvFormat = CSVFormat.DEFAULT .withAllowMissingColumnNames() .withTrim() .withQuoteMode(QuoteMode.MINIMAL) .withErrorOnMissingColumnNames(false); try (StringReader reader = new StringReader(wonkyCsv); CSVParser parser = new CSVParser(reader, csvFormat)) { List<CSVRecord> records = parser.getRecords(); for (CSVRecord record : records) { if (record.getRecordNumber() == 1) { // The first record is parsed as a single field—split it manually String fullField = record.get(0); int firstComma = fullField.indexOf(','); String id = fullField.substring(0, firstComma); String details = fullField.substring(firstComma + 1).replace("\"", ""); System.out.printf("Record 1: ID=%s, Details=%s%n", id, details); } else { // Normal records work as expected System.out.printf("Record %d: %s%n", record.getRecordNumber(), record); } } } catch (Exception e) { // Fallback if the parser still throws an error—handle the line manually System.out.printf("Hit a parsing snag: %s%n", e.getMessage()); } } }
The key here is that the lenient format lets the parser read the broken line as a single field, which we then split manually into the two columns you need.
Option 2: OpenCSV (Another Popular Choice)
OpenCSV is another solid library that makes handling messy CSVs straightforward. Here's how to use it:
Add the Maven dependency:
<dependency> <groupId>com.opencsv</groupId> <artifactId>opencsv</artifactId> <version>5.6</version> </dependency>
Code example:
import com.opencsv.CSVReader; import com.opencsv.exceptions.CsvValidationException; import java.io.StringReader; public class OpenCsvBrokenLineHandler { public static void main(String[] args) { String wonkyCsv = "1,\"Robert,Adams,Washington,US\n2,\"Madhu\",Grant,Oregon,US\n3,Mohan,Young,Texas,US"; try (CSVReader reader = new CSVReader(new StringReader(wonkyCsv))) { String[] nextLine; int lineCount = 0; while ((nextLine = reader.readNext()) != null) { lineCount++; if (lineCount == 1) { // Fix the broken first line String fullLine = nextLine[0]; int firstComma = fullLine.indexOf(','); String id = fullLine.substring(0, firstComma); String details = fullLine.substring(firstComma + 1).replace("\"", ""); System.out.printf("Line 1: ID=%s, Details=%s%n", id, details); } else { // Clean up quoted fields and print for (int i = 0; i < nextLine.length; i++) { nextLine[i] = nextLine[i].replace("\"", "").trim(); } System.out.printf("Line %d: %s%n", lineCount, String.join(", ", nextLine)); } } } catch (CsvValidationException e) { // Handle validation errors by targeting the problematic line System.out.printf("Validation error on line %d: %s%n", e.getLineNumber(), e.getMessage()); } catch (Exception e) { e.printStackTrace(); } } }
Option 3: Manual Parsing (No Dependencies)
If you don't want to add a library, you can read the CSV line-by-line and handle the broken record manually. Just note that this works best for simple CSV structures:
import java.io.BufferedReader; import java.io.StringReader; public class ManualCsvParser { public static void main(String[] args) { String wonkyCsv = "1,\"Robert,Adams,Washington,US\n2,\"Madhu\",Grant,Oregon,US\n3,Mohan,Young,Texas,US"; try (BufferedReader reader = new BufferedReader(new StringReader(wonkyCsv))) { String line; int lineCount = 0; while ((line = reader.readLine()) != null) { lineCount++; if (lineCount == 1) { // Split the broken line directly int firstComma = line.indexOf(','); String id = line.substring(0, firstComma); String details = line.substring(firstComma + 1).replace("\"", ""); System.out.printf("Line 1: ID=%s, Details=%s%n", id, details); } else { // Split normal lines while ignoring commas inside quotes String[] fields = line.split(",(?=(?:[^\"]*\"[^\"]*\")*[^\"]*$)"); for (int i = 0; i < fields.length; i++) { fields[i] = fields[i].replace("\"", "").trim(); } System.out.printf("Line %d: %s%n", lineCount, String.join(", ", fields)); } } } catch (Exception e) { e.printStackTrace(); } } }
The regex ,(?=(?:[^\"]*\"[^\"]*\")*[^\"]*$) ensures we don't split commas that are inside properly quoted fields.
Quick Recap
- For enterprise projects, stick with Apache Commons CSV or OpenCSV—they handle most edge cases out of the box and are well-maintained.
- Manual parsing works for simple scenarios but can get messy if your CSV has more complex edge cases (like escaped quotes).
- The core idea is to identify the broken line (either via parser exceptions or unexpected field counts) and manually split/clean it to get the columns you need.
内容的提问来源于stack exchange,提问作者user9212268

