使用Scanner解析CSV文件遇InputMismatchException及循环逻辑疑问
CSV解析InputMismatchException问题排查与解决
问题描述
有如下格式的CSV文件:
Company,Name,12,13,18 Company,Name,12,13,18 Company,Name,12,13,18 Company,Name,12,13,18 Company,Name,12,13,18
使用以下Java代码解析时,每行最后一个数值(如18)会触发InputMismatchException异常,同时对循环逻辑的正确性存在疑问:
public void readFile(String Path) { try { Scanner sc = new Scanner(Paths.get(Path)).useDelimiter(","); while (sc.hasNextLine()) { while (sc.hasNext()) { System.out.println(sc.next()); System.out.println(sc.next()); System.out.println(sc.nextInt()); System.out.println(sc.nextInt()); System.out.println(sc.nextInt()); } sc.close(); } } catch(IOException e) { System.out.println("Cannot find file: " + Path); } }
异常原因分析
- 分隔符设置问题:你仅将Scanner的分隔符设为逗号
,,但CSV每行最后一个元素后是换行符而非逗号。此时Scanner会把换行符和下一行的首个字段拼接成一个token,比如第一行的18会和下一行的Company拼成18\nCompany,调用nextInt()解析这个包含非数字字符的token时,必然抛出InputMismatchException。 - 循环结构错误:
- 内层
while(sc.hasNext())会一次性读取文件中所有token,外层while(sc.hasNextLine())只会执行一次,因为所有内容已被内层循环读完。 sc.close()放在外层循环内部,第一次循环就会关闭Scanner,后续循环无法继续执行。
- 内层
代码修正方案
方案1:调整分隔符,包含逗号和换行
将分隔符设置为匹配逗号或换行符,确保每个字段被正确拆分:
public void readFile(String Path) { try { // 匹配逗号、Windows换行符(\r\n)和Linux换行符(\n) Scanner sc = new Scanner(Paths.get(Path)).useDelimiter(",|\r\n|\n"); while (sc.hasNext()) { System.out.println(sc.next()); System.out.println(sc.next()); System.out.println(sc.nextInt()); System.out.println(sc.nextInt()); System.out.println(sc.nextInt()); } sc.close(); } catch(IOException e) { System.out.println("Cannot find file: " + Path); } }
方案2:逐行读取后拆分(更直观)
先读取整行,再拆分每行字段,彻底避免分隔符混淆问题:
public void readFile(String Path) { try { Scanner sc = new Scanner(Paths.get(Path)); while (sc.hasNextLine()) { String line = sc.nextLine(); String[] tokens = line.split(","); System.out.println(tokens[0]); System.out.println(tokens[1]); System.out.println(Integer.parseInt(tokens[2])); System.out.println(Integer.parseInt(tokens[3])); System.out.println(Integer.parseInt(tokens[4])); } sc.close(); } catch(IOException e) { System.out.println("Cannot find file: " + Path); } catch(NumberFormatException e) { System.out.println("数字格式错误"); } }
关于循环逻辑的理解
你的理解不正确:
- 外层
while(sc.hasNextLine())是判断是否还有未读取的行,但内层while(sc.hasNext())会基于设置的分隔符读取整个文件的token,而非当前行的token,因此会一次性读完所有内容,外层循环仅执行一次。 - 如果要实现“外层循环遍历每行,内层循环遍历该行的每个token”,应该先读取整行,再对该行内容创建单独的Scanner或拆分字符串,如方案2的实现方式。
内容的提问来源于stack exchange,提问作者Nathan
相关产品推荐
相关产品推荐

