哈希计算内存耗尽且速度持续变慢问题排查与解决诉求
解决重复计算大文件哈希时速度持续变慢的问题
问题背景
开发的JavaFX GUI哈希工具在重复计算1GB文件哈希时出现性能衰减:首次计算耗时约2秒,后续重复计算耗时逐步增至76秒。简化测试代码(核心逻辑为一次性读取文件为字节数组再计算MD5)重复执行5次后,出现明显的速度下降,怀疑是内存占用过高导致,需要优化内存使用,让每次计算都能保持首次的速度。
原测试代码
package main; import java.io.File; import java.io.FileInputStream; import java.io.IOException; import java.security.MessageDigest; import java.security.NoSuchAlgorithmException; import org.springframework.util.StopWatch; public class Main { public static void main(String[] args) throws NoSuchAlgorithmException { // 1GB Test file StopWatch sw = new StopWatch(); // Hash the file 5 times and measure the time for each hashing for (int i = 0; i < 5; i++) { sw.start(); String hash = encrypt(inputparser("C:\\Users\\Thend\\Downloads\\1GB.bin")); sw.stop(); // Print execution time at each cycle System.out.println("Execution time: "+String.format("%.5f", sw.getTotalTimeMillis() / 1000.0f)+"sec"+" Hash: "+hash); } } public static byte[] inputparser(String path){ File f = new File(path); byte[] bytes = new byte[(int) f.length()]; FileInputStream fis = null; try { fis = new FileInputStream(f); // read file into bytes[] fis.read(bytes); if (fis != null) { fis.close(); } } catch (IOException e) { e.printStackTrace(); } return bytes; } public static String encrypt(byte[] bytes) throws NoSuchAlgorithmException { MessageDigest md = MessageDigest.getInstance("MD5"); StringBuilder sb = new StringBuilder(); md.reset(); md.update(bytes); byte[] hashed_bytes = md.digest(); // Convert bytes[] (in decimal format) to hexadecimal for (int i = 0; i < hashed_bytes.length; i++) { sb.append(Integer.toString((hashed_bytes[i] & 0xff) + 0x100, 16).substring(1)); } // Return hashed String in hex format String hashedByteArray = sb.toString(); return hashedByteArray; } }
原控制台输出(计时存在错误,显示为累计时间)
"C:\Program Files\Java\jdk-18.0.2\bin\java.exe" "-javaagent:C:\Program Files\JetBrains\IntelliJ IDEA Community Edition 2022.2\lib\idea_rt.jar=54323:C:\Program Files\JetBrains\IntelliJ IDEA Community Edition 2022.2\bin" -Dfile.encoding=UTF-8 -Dsun.stdout.encoding=UTF-8 -Dsun.stderr.encoding=UTF-8 -classpath C:\Users\Thend\intellij-workspace\HashTimeTest\out\production\HashTimeTest;C:\Users\Thend\Downloads\springframework-5.1.0.jar main.Main Execution time: 2,04100sec Hash: e5c834fbdaa6bfd8eac5eb9404eefdd4 Execution time: 3,70900sec Hash: e5c834fbdaa6bfd8eac5eb9404eefdd4 Execution time: 5,42100sec Hash: e5c834fbdaa6bfd8eac5eb9404eefdd4 Execution time: 7,09600sec Hash: e5c834fbdaa6bfd8eac5eb9404eefdd4 Execution time: 8,75500sec Hash: e5c834fbdaa6bfd8eac5eb9404eefdd4 Process finished with exit code 0
问题分析
- 内存占用过高:每次循环都一次性将1GB文件加载到byte数组中,这些大数组会占用大量堆内存,且循环结束后不会被立即回收(JVM的GC触发时机不确定),导致后续计算只能使用老年代内存,甚至触发频繁的Full GC,严重拖慢速度。
- 计时错误:原代码使用
sw.getTotalTimeMillis()获取的是累计执行时间,而非单次计算的耗时,导致对性能衰减的判断存在偏差,但实际重复执行时的内存压力确实会导致真实的性能下降。
解决方案
优化点
- 改用流式分块读取文件,避免一次性加载整个文件到内存,从根源减少内存占用。
- 修正StopWatch的使用方式,正确统计单次计算的耗时。
- 显式解除大对象引用,帮助JVM及时回收内存。
- 优化哈希计算的字节转十六进制逻辑,提升效率。
修改后的代码
package main; import java.io.BufferedInputStream; import java.io.File; import java.io.FileInputStream; import java.io.IOException; import java.security.MessageDigest; import java.security.NoSuchAlgorithmException; import org.springframework.util.StopWatch; public class Main { private static final int BUFFER_SIZE = 8192; // 8KB缓冲,平衡IO和内存 public static void main(String[] args) throws NoSuchAlgorithmException { StopWatch sw = new StopWatch(); String filePath = "C:\\Users\\Thend\\Downloads\\1GB.bin"; for (int i = 0; i < 5; i++) { sw.start("Hash calculation " + (i+1)); String hash = encrypt(filePath); sw.stop(); // 输出单次任务的耗时,而非累计时间 System.out.printf("Execution time: %.5fsec Hash: %s%n", sw.getLastTaskTimeMillis() / 1000.0f, hash); // 显式触发GC(测试场景用,生产环境不建议依赖) System.gc(); System.runFinalization(); } } public static String encrypt(String filePath) throws NoSuchAlgorithmException { MessageDigest md = MessageDigest.getInstance("MD5"); try (BufferedInputStream bis = new BufferedInputStream(new FileInputStream(filePath))) { byte[] buffer = new byte[BUFFER_SIZE]; int bytesRead; while ((bytesRead = bis.read(buffer)) != -1) { md.update(buffer, 0, bytesRead); } } catch (IOException e) { e.printStackTrace(); return null; } byte[] hashedBytes = md.digest(); // 优化十六进制转换,使用更快的实现 StringBuilder sb = new StringBuilder(hashedBytes.length * 2); for (byte b : hashedBytes) { sb.append(String.format("%02x", b)); } return sb.toString(); } }
优化说明
- 流式读取:使用
BufferedInputStream分块读取文件,每次仅加载8KB缓冲到内存,彻底避免1GB大数组的内存占用,无论重复执行多少次,内存消耗都保持在极低水平。 - 正确计时:使用
sw.getLastTaskTimeMillis()获取单次任务的耗时,准确反映每次计算的真实速度。 - 内存回收:循环末尾调用
System.gc()(测试场景下),帮助JVM及时回收临时对象,确保每次计算的内存环境一致。 - 转换优化:用
String.format("%02x", b)替代原有的复杂整数转换逻辑,代码更简洁且效率更高。
内容的提问来源于stack exchange,提问作者Thend
相关产品推荐
相关产品推荐

