使用BufferedReader读取文件时Integer.parseInt()仅metro.txt抛NumberFormatException
问题描述
尝试用Java读取文本文件中的数字,遇到如下异常:
- 作业提供的
metro.txt前几行内容:
376 933 0000 Abbesses 0001 Alexandre Dumas
- 自制
testnum.txt前几行内容:
444 555 6666 flowers 8888 pumpkin patch
使用以下Java代码读取:
import java.io.BufferedReader; import java.io.FileNotFoundException; import java.io.FileReader; import java.io.IOException; public class test { public static void main(String[] args) { try { FileReader input = new FileReader("metro.txt"); BufferedReader reader = new BufferedReader(input); String raw_line = reader.readLine(); String[] split_line = raw_line.split(" ", 2); int num_vertices = Integer.parseInt(split_line[0]); int num_edges = Integer.parseInt(split_line[1]); } catch (FileNotFoundException ex) { // catch filereader System.out.println(ex); } catch (IOException ex) { // catch bufferedreader System.out.println(ex); } } }
读取metro.txt时抛出异常:
Exception in thread "main" java.lang.NumberFormatException: For input string: "376" at java.base/java.lang.NumberFormatException.forInputString(NumberFormatException.java:67) at java.base/java.lang.Integer.parseInt(Integer.java:668) at java.base/java.lang.Integer.parseInt(Integer.java:786) at test.main(test.java:15)
读取testnum.txt无异常,且split_line[0]与字符串"376"比较返回false,怀疑是文件编码导致问题,请求排查根源及解决方法。
问题根源
核心原因是**metro.txt带有UTF-8 BOM(字节顺序标记)**。作业提供的文件通常是Windows环境下生成的UTF-8文件,会在文件开头添加不可见的BOM字符(\uFEFF):
FileReader默认使用系统默认编码读取文件,当文件是带BOM的UTF-8时,BOM字符会被读入字符串开头,导致split后的split_line[0]实际是"\uFEFF376",而非肉眼看到的"376"。- 这个不可见字符无法通过普通字符串对比察觉,但会导致
Integer.parseInt()解析失败,同时直接和"376"比较返回false。
解决方法
方法1:指定编码读取并跳过BOM
用InputStreamReader指定UTF-8编码读取,手动判断并跳过BOM字符:
import java.io.BufferedReader; import java.io.FileInputStream; import java.io.IOException; import java.io.InputStreamReader; public class test { public static void main(String[] args) { try (FileInputStream fis = new FileInputStream("metro.txt"); InputStreamReader isr = new InputStreamReader(fis, "UTF-8"); BufferedReader reader = new BufferedReader(isr)) { String raw_line = reader.readLine(); // 去除开头的BOM字符 if (raw_line != null && raw_line.startsWith("\uFEFF")) { raw_line = raw_line.substring(1); } String[] split_line = raw_line.split(" ", 2); int num_vertices = Integer.parseInt(split_line[0]); int num_edges = Integer.parseInt(split_line[1]); System.out.println("顶点数:" + num_vertices + ",边数:" + num_edges); } catch (IOException ex) { System.out.println(ex); } } }
方法2:读取后清理非有效字符
如果不确定文件编码,可直接对读取的字符串做清理,只保留数字和空白字符:
String raw_line = reader.readLine(); if (raw_line != null) { // 替换所有非数字和空白的字符为空 raw_line = raw_line.replaceAll("[^0-9\\s]", ""); String[] split_line = raw_line.split(" ", 2); // 后续解析逻辑不变 }
方法3:使用Java NIO读取(更简洁)
用Files.readAllLines指定UTF-8编码读取,再处理BOM:
import java.nio.file.Files; import java.nio.file.Paths; import java.util.List; public class test { public static void main(String[] args) { try { List<String> lines = Files.readAllLines(Paths.get("metro.txt"), java.nio.charset.StandardCharsets.UTF_8); String raw_line = lines.get(0); if (raw_line.startsWith("\uFEFF")) { raw_line = raw_line.substring(1); } String[] split_line = raw_line.split(" ", 2); int num_vertices = Integer.parseInt(split_line[0]); int num_edges = Integer.parseInt(split_line[1]); System.out.println("顶点数:" + num_vertices + ",边数:" + num_edges); } catch (IOException ex) { System.out.println(ex); } } }
内容的提问来源于stack exchange,提问作者MatRanc
相关产品推荐
相关产品推荐

