Java读取CSV文件报错,需自动修复下载后文件格式问题
解决Apache Commons CSV读取末尾带分号的CSV文件报错问题
问题根源
你遇到的错误是因为原始CSV文件的表头行末尾多了一个分号,导致Apache Commons CSV解析时识别出一个空的表头名称,而默认配置不允许空表头,因此抛出IllegalArgumentException。手动保存后,这个末尾分号被移除,所以解析正常。
解决方案
方案1:修改CSVFormat配置,允许空表头名称
直接在CSVFormat中添加配置,允许解析包含空表头的文件,同时跳过空表头的输出:
package collectData; import java.io.FileReader; import java.io.IOException; import java.io.Reader; import java.nio.charset.Charset; import org.apache.commons.csv.CSVFormat; import org.apache.commons.csv.CSVParser; import org.apache.commons.csv.CSVRecord; public class testCSV { public static void main(String[] args) throws IOException { String csvFilePath = "csv_files\\20240403_OV64E_A.csv"; Charset encoding = Charset.forName("ISO-8859-1"); try (Reader in = new FileReader(csvFilePath, encoding); CSVParser parser = CSVFormat.DEFAULT .withHeader() .withDelimiter(';') .withAllowMissingColumnNames() // 允许空表头 .parse(in)) { for (CSVRecord record : parser) { for (String headerName : parser.getHeaderNames()) { if (!headerName.isEmpty()) { // 跳过空表头,避免无效输出 System.out.print(headerName + ": " + record.get(headerName) +" "); } } System.out.println(); } } } }
方案2:预处理文件,移除表头行末尾的分号
如果不想允许空表头,可以先读取文件内容,预处理表头行去掉末尾的分号,再传递给CSVParser:
package collectData; import java.io.BufferedReader; import java.io.FileReader; import java.io.IOException; import java.io.StringReader; import java.nio.charset.Charset; import java.util.stream.Collectors; import org.apache.commons.csv.CSVFormat; import org.apache.commons.csv.CSVParser; import org.apache.commons.csv.CSVRecord; public class testCSV { public static void main(String[] args) throws IOException { String csvFilePath = "csv_files\\20240403_OV64E_A.csv"; Charset encoding = Charset.forName("ISO-8859-1"); try (BufferedReader br = new BufferedReader(new FileReader(csvFilePath, encoding))) { String firstLine = br.readLine(); // 移除表头行末尾的分号(如果存在) if (firstLine != null && firstLine.endsWith(";")) { firstLine = firstLine.substring(0, firstLine.length() - 1); } // 拼接处理后的表头和剩余内容 String processedContent = firstLine + "\n" + br.lines().collect(Collectors.joining("\n")); try (Reader in = new StringReader(processedContent); CSVParser parser = CSVFormat.DEFAULT .withHeader() .withDelimiter(';') .parse(in)) { for (CSVRecord record : parser) { for (String headerName : parser.getHeaderNames()) { System.out.print(headerName + ": " + record.get(headerName) +" "); } System.out.println(); } } } } }
补充说明
- 数据行的空字段(比如
003298;4.1;Figée;;;;;)会被Apache Commons CSV自动处理,返回空字符串,无需额外配置。 - 关于
003298变成3298的问题:这是手动保存时Excel自动格式化导致的,CSVParser解析时会保留原始字符串,不会丢失前导零。
内容的提问来源于stack exchange,提问作者Mouad Thf
相关产品推荐
相关产品推荐

