如何用Java 8按表头读取CSV文件中的两列数据
问题
需要通过表头col5和col7读取CSV文件中的两列数据,现有代码仅能读取单列(col5),希望将这两列数据存入同一个Map中(选择Map是因为后续需要对数据进行处理)。
现有代码
public static List<String> getData() throws IOException { String titleToSearchFor = "col5"; Path path = Paths.get("D:/blackduck_sample.txt"); if (Files.exists(path)) { List<String> lines = Files.readAllLines(path); List<String> columns = Arrays.asList(lines.get(0).split(",")); int titleIndex = columns.indexOf(titleToSearchFor); List<String> values = lines.stream().skip(1).map(line -> Arrays.asList(line.split(","))) .map(list -> list.get(titleIndex)).filter(Objects::nonNull).filter(s -> s.trim().length() > 0) .collect(Collectors.toList()); return values; } return new ArrayList<>(); }
样本CSV文件
col1,col2,col3,col4,col5,col6,col7,col8 CVM Finding,12345,sdc,CAP,12345:sdc-posistion-svc:SnakeYAML:2.1:3,TRUE,FALSE,OverDue CVM Finding,12345,sdc,CAP,12345:sdc-event-generation-svc:SpringBoot:2.1.2.Release:1,TRUE,FALSE,Overdue CVM Finding,12345,sdc,CAP,12345:sdc-event-subscription-svc:Apache Tomcat:9.0.70:1,TRUE,FALSE,Due in <= 90 CVM Finding,12345,sdc,CAP,12345:sdc-bh-registry-svc:Jettison - Json Stax implementation:1.5.3:1,TRUE,FALSE,Due in <=90 CVM Finding,12345,sdc,CAP,12345:sdc-event-generation-svc:Apache Tomcat:9.0.70:2,TRUE,FALSE,Due in >90 CVM Finding,12345,sdc,CAP,12345:sdc-posistion-svc:jackson-databind:2.14:1,TRUE,FALSE,OverDue CVM Finding,12345,sdc,CAP,12345:sdc-event-subscription-svc:jackson-databind:2.14:1,TRUE,FALSE,Due in <= 90 CVM Finding,12345,sdc,CAP,12345:sdc-event-subscription-svc:SpringBoot:2.1.2.Release:1,TRUE,FALSE,Overdue CVM Finding,12345,sdc,CAP,12345:sdc-posistion-svc:SpringBoot:2.1.2.Release:1,TRUE,FALSE,OverDue
之前找到的参考方法(仅按索引读取)
String[] cols = line.split(cvsSplitBy); System.out.println("Column 5= " + cols[5] + " , Column 7=" + cols[7]);
期望输出
12345:sdc-posistion-svc:SnakeYAML:2.1:3,OverDue 12345:sdc-event-generation-svc:SpringBoot:2.1.2.Release:1,Overdue 12345:sdc-event-subscription-svc:Apache Tomcat:9.0.70:1,Due in <= 90 12345:sdc-bh-registry-svc:Jettison - Json Stax implementation:1.5.3:1,Due in <=90 12345:sdc-event-generation-svc:Apache Tomcat:9.0.70:2,Due in >90 12345:sdc-posistion-svc:jackson-databind:2.14:1,OverDue 12345:sdc-event-subscription-svc:jackson-databind:2.14:1,Due in <= 90 12345:sdc-event-subscription-svc:SpringBoot:2.1.2.Release:1,Overdue 12345:sdc-posistion-svc:SpringBoot:2.1.2.Release:1,OverDue
解决方案
完全可以将两列数据存入同一个Map中,以下提供两种常见实现方式:
方式1:将col5作为Key,目标列作为Value存入Map
(注:你的期望输出第二列对应样本CSV的col8,若需匹配期望输出,只需将代码中的col7Title改为col8即可)
import java.io.IOException; import java.nio.file.Files; import java.nio.file.Paths; import java.util.Map; import java.util.Objects; import java.util.stream.Collectors; public class CSVReader { public static Map<String, String> getDataAsMap() throws IOException { String col5Title = "col5"; String targetColTitle = "col7"; // 若匹配期望输出,改为"col8" Path path = Paths.get("D:/blackduck_sample.txt"); if (Files.exists(path)) { var lines = Files.readAllLines(path); if (lines.isEmpty()) return Map.of(); // 获取表头索引 var columns = java.util.Arrays.asList(lines.get(0).split(",")); int col5Index = columns.indexOf(col5Title); int targetColIndex = columns.indexOf(targetColTitle); // 校验表头是否存在 if (col5Index == -1 || targetColIndex == -1) { throw new IllegalArgumentException("指定表头不存在:" + col5Title + " 或 " + targetColTitle); } // 读取数据并封装为Map return lines.stream() .skip(1) .map(line -> line.split(",")) .filter(parts -> parts.length > Math.max(col5Index, targetColIndex)) .map(parts -> { String col5Val = parts[col5Index].trim(); String targetVal = parts[targetColIndex].trim(); return Map.entry(col5Val, targetVal); }) .filter(entry -> !entry.getKey().isEmpty() && !entry.getValue().isEmpty()) .collect(Collectors.toMap( Map.Entry::getKey, Map.Entry::getValue, (existing, replacement) -> existing // 重复Key保留第一个值,可按需修改 )); } return Map.of(); } // 生成符合期望的字符串格式输出 public static java.util.List<String> getFormattedOutput() throws IOException { var dataMap = getDataAsMap(); return dataMap.entrySet().stream() .map(entry -> entry.getKey() + "," + entry.getValue()) .collect(Collectors.toList()); } public static void main(String[] args) throws IOException { getFormattedOutput().forEach(System.out::println); } }
方式2:每行数据封装为Map(键为表头名,值为对应列数据)
适合需要保留表头与值对应关系的场景:
public static java.util.List<Map<String, String>> getRowData() throws IOException { String col5Title = "col5"; String targetColTitle = "col7"; // 若匹配期望输出,改为"col8" Path path = Paths.get("D:/blackduck_sample.txt"); if (Files.exists(path)) { var lines = Files.readAllLines(path); if (lines.isEmpty()) return new java.util.ArrayList<>(); var columns = java.util.Arrays.asList(lines.get(0).split(",")); int col5Index = columns.indexOf(col5Title); int targetColIndex = columns.indexOf(targetColTitle); if (col5Index == -1 || targetColIndex == -1) { throw new IllegalArgumentException("指定表头不存在:" + col5Title + " 或 " + targetColTitle); } return lines.stream() .skip(1) .map(line -> line.split(",")) .filter(parts -> parts.length > Math.max(col5Index, targetColIndex)) .map(parts -> { String col5Val = parts[col5Index].trim(); String targetVal = parts[targetColIndex].trim(); return Map.of(col5Title, col5Val, targetColTitle, targetVal); }) .filter(map -> !map.get(col5Title).isEmpty() && !map.get(targetColTitle).isEmpty()) .collect(Collectors.toList()); } return new java.util.ArrayList<>(); }
代码说明
- 避免硬编码索引,通过表头名动态获取列位置,适配表头变动场景
- 增加数据校验,防止数组越界和空值干扰
- 提供两种Map存储方式,可根据后续处理需求选择
内容的提问来源于stack exchange,提问作者Selenium New bie
相关产品推荐
相关产品推荐

