WebLogic下用Scanner逐行读文件生成CensusData数组触发Java堆内存溢出
堆内存溢出原因分析
- 你所有创建的
CensusData实例都被ArrayList持有强引用,GC无法回收,全量数据本身就会占用堆内存,并不是没在循环内声明局部变量就不会产生堆内存占用。 - 代码中存在大量冗余操作放大了内存消耗:每行数据重复调用了41次
s.split(","),每次调用都会生成新的字符串数组,同时产生大量临时字符串对象,这部分额外内存占用甚至可能超过业务数据本身的占用量。 ArrayList初始容量仅设置为43,当数据量超过容量时会触发1.5倍扩容,扩容过程中需要创建新的数组并复制旧数据,会产生临时的双倍内存占用峰值,很容易触发OOM。- Java 8及更早版本中,
split、trim生成的子字符串会持有原整行字符串的引用,导致整行的大字符串对象无法被GC回收,进一步放大内存占用。 Scanner本身基于正则匹配实现,内部缓冲机制也会产生额外的内存开销。
可行优化方案
所有方案均不需要调整JVM堆内存参数,且能保证获取完整的CensusData数组:
1. 基础优化(改造成本最低,可覆盖绝大多数场景)
核心是消除冗余操作,降低不必要的内存开销,优化后代码如下:
public static CensusData[] load(String path) throws Exception { // 预估算100M CSV文件大概有10万行,提前初始化容量避免扩容开销,可根据实际行数调整 List<CensusData> list = new ArrayList<>(100000); // 用更轻量的BufferedReader替代Scanner,减少额外内存占用 try (BufferedReader br = new BufferedReader(new InputStreamReader(new FileInputStream(path), "UTF-8"))) { String s; while ((s = br.readLine()) != null) { // 每行只split一次,复用拆分结果 String[] fields = s.split(","); // 对于Java 8及更早版本,字符串字段用new String包裹,切断和原行字符串的引用 list.add(new CensusData( Integer.parseInt(fields[0].trim()), new String(fields[1].trim()), Integer.parseInt(fields[2].trim()), Integer.parseInt(fields[3].trim()), new String(fields[4].trim()), Integer.parseInt(fields[5].trim()), new String(fields[6].trim()), new String(fields[7].trim()), new String(fields[8].trim()), new String(fields[9].trim()), new String(fields[10].trim()), new String(fields[11].trim()), new String(fields[12].trim()), new String(fields[13].trim()), new String(fields[14].trim()), new String(fields[15].trim()), Integer.parseInt(fields[16].trim()), Integer.parseInt(fields[17].trim()), Integer.parseInt(fields[18].trim()), new String(fields[19].trim()), new String(fields[20].trim()), new String(fields[21].trim()), new String(fields[22].trim()), new String(fields[24].trim()), new String(fields[25].trim()), new String(fields[26].trim()), new String(fields[27].trim()), new String(fields[28].trim()), new String(fields[29].trim()), new String(fields[30].trim()), new String(fields[31].trim()), new String(fields[32].trim()), new String(fields[33].trim()), new String(fields[34].trim()), new String(fields[35].trim()), new String(fields[36].trim()), new String(fields[37].trim()), new String(fields[38].trim()), Integer.parseInt(fields[39].trim()), new String(fields[40].trim()) )); } } catch (Exception e) { e.printStackTrace(); throw e; } return list.toArray(new CensusData[0]); }
2. 进阶优化(内存占用进一步降低30%以上)
如果基础优化后依然内存不足,可以采用享元模式复用重复字符串:如果你的CensusData中存在大量重复内容的字段(如地域、类别等枚举类字段),可以对这些字段做缓存复用,避免重复创建相同内容的字符串对象:
// 全局缓存,只缓存出现过的字符串,重复内容直接复用 private static final Map<String, String> STRING_CACHE = new HashMap<>(); // 构造CensusData时,字符串字段替换为缓存复用的逻辑 private static String getCachedString(String raw) { String trim = raw.trim(); return STRING_CACHE.computeIfAbsent(trim, k -> new String(trim)); }
构造CensusData时所有字符串字段都用getCachedString(fields[x])替代原来的trim逻辑即可。
内容的提问来源于stack exchange,提问作者reza
相关产品推荐
相关产品推荐

