You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

哈希计算内存耗尽且速度持续变慢问题排查与解决诉求

解决重复计算大文件哈希时速度持续变慢的问题

问题背景

开发的JavaFX GUI哈希工具在重复计算1GB文件哈希时出现性能衰减:首次计算耗时约2秒,后续重复计算耗时逐步增至76秒。简化测试代码(核心逻辑为一次性读取文件为字节数组再计算MD5)重复执行5次后,出现明显的速度下降,怀疑是内存占用过高导致,需要优化内存使用,让每次计算都能保持首次的速度。

原测试代码

package main;

import java.io.File;
import java.io.FileInputStream;
import java.io.IOException;
import java.security.MessageDigest;
import java.security.NoSuchAlgorithmException;
import org.springframework.util.StopWatch;

public class Main {

    public static void main(String[] args) throws NoSuchAlgorithmException {

        // 1GB Test file
        StopWatch sw = new StopWatch();
        // Hash the file 5 times and measure the time for each hashing
        for (int i = 0; i < 5; i++) {
            sw.start();
                String hash = encrypt(inputparser("C:\\Users\\Thend\\Downloads\\1GB.bin"));
            sw.stop();
            // Print execution time at each cycle
            System.out.println("Execution time: "+String.format("%.5f", sw.getTotalTimeMillis() / 1000.0f)+"sec"+" Hash: "+hash);
        }
    }

    public static byte[] inputparser(String path){
        File f = new File(path);
        byte[] bytes = new byte[(int) f.length()];
        FileInputStream fis = null;
        try {
            fis = new FileInputStream(f);
            // read file into bytes[]
            fis.read(bytes);
            if (fis != null) {
                fis.close();
            }
        } catch (IOException e) {
            e.printStackTrace();
        }
        return bytes;
    }

    public static String encrypt(byte[] bytes) throws NoSuchAlgorithmException {
        MessageDigest md = MessageDigest.getInstance("MD5");
        StringBuilder sb = new StringBuilder();

        md.reset();
        md.update(bytes);

        byte[] hashed_bytes = md.digest();

        // Convert bytes[] (in decimal format) to hexadecimal
        for (int i = 0; i < hashed_bytes.length; i++) {
            sb.append(Integer.toString((hashed_bytes[i] & 0xff) + 0x100, 16).substring(1));
        }

        // Return hashed String in hex format
        String hashedByteArray = sb.toString();
        return hashedByteArray;
    }
}

原控制台输出(计时存在错误,显示为累计时间)

"C:\Program Files\Java\jdk-18.0.2\bin\java.exe" "-javaagent:C:\Program Files\JetBrains\IntelliJ IDEA Community Edition 2022.2\lib\idea_rt.jar=54323:C:\Program Files\JetBrains\IntelliJ IDEA Community Edition 2022.2\bin" -Dfile.encoding=UTF-8 -Dsun.stdout.encoding=UTF-8 -Dsun.stderr.encoding=UTF-8 -classpath C:\Users\Thend\intellij-workspace\HashTimeTest\out\production\HashTimeTest;C:\Users\Thend\Downloads\springframework-5.1.0.jar main.Main
Execution time: 2,04100sec Hash: e5c834fbdaa6bfd8eac5eb9404eefdd4
Execution time: 3,70900sec Hash: e5c834fbdaa6bfd8eac5eb9404eefdd4
Execution time: 5,42100sec Hash: e5c834fbdaa6bfd8eac5eb9404eefdd4
Execution time: 7,09600sec Hash: e5c834fbdaa6bfd8eac5eb9404eefdd4
Execution time: 8,75500sec Hash: e5c834fbdaa6bfd8eac5eb9404eefdd4

Process finished with exit code 0

问题分析

  • 内存占用过高:每次循环都一次性将1GB文件加载到byte数组中,这些大数组会占用大量堆内存,且循环结束后不会被立即回收(JVM的GC触发时机不确定),导致后续计算只能使用老年代内存,甚至触发频繁的Full GC,严重拖慢速度。
  • 计时错误:原代码使用sw.getTotalTimeMillis()获取的是累计执行时间,而非单次计算的耗时,导致对性能衰减的判断存在偏差,但实际重复执行时的内存压力确实会导致真实的性能下降。

解决方案

优化点

  • 改用流式分块读取文件,避免一次性加载整个文件到内存,从根源减少内存占用。
  • 修正StopWatch的使用方式,正确统计单次计算的耗时。
  • 显式解除大对象引用,帮助JVM及时回收内存。
  • 优化哈希计算的字节转十六进制逻辑,提升效率。

修改后的代码

package main;

import java.io.BufferedInputStream;
import java.io.File;
import java.io.FileInputStream;
import java.io.IOException;
import java.security.MessageDigest;
import java.security.NoSuchAlgorithmException;
import org.springframework.util.StopWatch;

public class Main {

    private static final int BUFFER_SIZE = 8192; // 8KB缓冲,平衡IO和内存

    public static void main(String[] args) throws NoSuchAlgorithmException {
        StopWatch sw = new StopWatch();
        String filePath = "C:\\Users\\Thend\\Downloads\\1GB.bin";

        for (int i = 0; i < 5; i++) {
            sw.start("Hash calculation " + (i+1));
            String hash = encrypt(filePath);
            sw.stop();
            // 输出单次任务的耗时,而非累计时间
            System.out.printf("Execution time: %.5fsec Hash: %s%n", 
                sw.getLastTaskTimeMillis() / 1000.0f, hash);
            
            // 显式触发GC(测试场景用,生产环境不建议依赖)
            System.gc();
            System.runFinalization();
        }
    }

    public static String encrypt(String filePath) throws NoSuchAlgorithmException {
        MessageDigest md = MessageDigest.getInstance("MD5");
        try (BufferedInputStream bis = new BufferedInputStream(new FileInputStream(filePath))) {
            byte[] buffer = new byte[BUFFER_SIZE];
            int bytesRead;
            while ((bytesRead = bis.read(buffer)) != -1) {
                md.update(buffer, 0, bytesRead);
            }
        } catch (IOException e) {
            e.printStackTrace();
            return null;
        }

        byte[] hashedBytes = md.digest();
        // 优化十六进制转换,使用更快的实现
        StringBuilder sb = new StringBuilder(hashedBytes.length * 2);
        for (byte b : hashedBytes) {
            sb.append(String.format("%02x", b));
        }
        return sb.toString();
    }
}

优化说明

  1. 流式读取:使用BufferedInputStream分块读取文件,每次仅加载8KB缓冲到内存,彻底避免1GB大数组的内存占用,无论重复执行多少次,内存消耗都保持在极低水平。
  2. 正确计时:使用sw.getLastTaskTimeMillis()获取单次任务的耗时,准确反映每次计算的真实速度。
  3. 内存回收:循环末尾调用System.gc()(测试场景下),帮助JVM及时回收临时对象,确保每次计算的内存环境一致。
  4. 转换优化:用String.format("%02x", b)替代原有的复杂整数转换逻辑,代码更简洁且效率更高。

内容的提问来源于stack exchange,提问作者Thend

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.20 04:20:30