You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Java并行有序字符串构建效率不及串行,矩阵打印遇性能瓶颈

并行打印矩阵到CLI的性能瓶颈解决

我有一个在while{}循环中持续更新的int[][]矩阵,每次迭代后需要将矩阵新值打印到CLI。为配合程序其他部分的并行逻辑,我实现了并行打印,但要求打印顺序严格对应矩阵行顺序。当前实现的性能比串行打印还慢,怀疑CountDownLatch的使用是瓶颈。

原实现代码

主循环代码

int numbOfBlocks = Runtime.getRuntime().availableProcessors();
int blockSize = matrix.length/numbOfBlocks;
int endRow;

while(true) {
   // print matrix to CLI

   StringBuilder stringBuilder = new StringBuilder();
   CountDownLatch stringBuilderLatch = new CountDownLatch(numbOfBlocks);

   for (int i = 0; i < numbOfBlocks; i++) {
      endRow = (i == numbOfBlocks - 1) ? rows : (i+1) * blockSize;
      StringBuilderThread stringBuilderThread =
            new StringBuilderThread(matrix, i*blockSize, endRow, stringBuilderLatch, stringBuilder, i, numbOfBlocks);
      executor.execute(stringBuilderThread);
   }
   stringBuilderLatch.await();

   stringBuilder.append("___");
   System.out.println(stringBuilder.toString());
}

StringBuilderThread代码

import java.util.concurrent.CountDownLatch;

public class StringBuilderThread implements Runnable {

    private int[][] matrixUpd;
    private int startRow;
    private int endRow;
    private CountDownLatch latch;
    private StringBuilder mainStringBuilder;
    private int order;
    private int numbOfBlocks;

    public StringBuilderThread(int[][] matrixUpd,
                               int startRow,
                               int endRow,
                               CountDownLatch latch,
                               StringBuilder mainStringBuilder,
                               int order,
                               int numbOfBlocks) {
        this.matrixUpd = matrixUpd;
        this.startRow = startRow;
        this.endRow = endRow;
        this.latch = latch;
        this.mainStringBuilder = mainStringBuilder;
        this.order = order;
        this.numbOfBlocks = numbOfBlocks;
    }
    @Override
    public void run() {
        StringBuilder tempStringBuilder = new StringBuilder();
        int cols = matrixUpd[0].length;
        for (int i = startRow; i < endRow; i++) {
            tempStringBuilder.append("|");
            for (int j = 0; j < cols; j++) {
                if(matrixUpd[i][j] == 1) {
                    tempStringBuilder.append("*  ");
                } else {
                    tempStringBuilder.append(".  ");
                }
            }
            tempStringBuilder.append("|\n");
        }
        tempStringBuilder.append(" ");

        // order synchronization, so that threads write to stringbuilders in correct order
        long latchCount = latch.getCount();
        while (!(numbOfBlocks-latchCount == order)) {
            latchCount = latch.getCount();
        }
        mainStringBuilder.append(tempStringBuilder.toString());
        latch.countDown();
    }
}

原实现的核心问题

原代码的性能瓶颈集中在自旋等待+共享StringBuilder的同步逻辑:

  • 线程构建完子字符串后,通过持续轮询CountDownLatch计数等待写入顺序,属于无意义的忙等待,占用大量CPU资源。
  • 虽然线程用临时StringBuilder构建局部内容,但最终写入共享mainStringBuilder时,自旋逻辑本质上强制了串行写入,完全抵消了并行构建的优势,还额外增加了线程调度和轮询的开销。

优化方案与测试结果

针对问题测试了三种解决方案,覆盖不同并行实现思路:

测试方案

  • 串行打印:单线程遍历矩阵构建字符串并打印
  • Callable并行构建:每个线程构建对应块的字符串,通过Future收集结果后按顺序拼接
  • 数组并行流构建:利用Java 8+的并行流,将矩阵行分块并行处理,再按顺序合并结果

测试场景

测试了四种不同维度的矩阵及对应迭代次数:

  • 200x200,迭代100次
  • 1000x1000,迭代100次
  • 2000x2000,迭代200次
  • 4000x4000,迭代200次

最优方案:数组并行流实现

并行流方案在所有测试场景中性能最优,且代码实现最简洁,无需手动管理线程和同步逻辑。示例代码如下:

while(true) {
    // 并行处理矩阵行,构建每行的字符串,再按顺序拼接
    String matrixStr = IntStream.range(0, rows)
            .parallel()
            .mapToObj(i -> {
                StringBuilder rowBuilder = new StringBuilder("|");
                for (int j = 0; j < cols; j++) {
                    rowBuilder.append(matrix[i][j] == 1 ? "*  " : ".  ");
                }
                return rowBuilder.append("|\n").toString();
            })
            .collect(Collectors.joining())
            + "___";
    
    System.out.println(matrixStr);
}

方案优势

  • 并行流自动管理线程池(默认使用ForkJoinPool.commonPool()),避免手动线程调度的开销
  • 无需额外同步组件(如CountDownLatch),通过流的collect操作自然保证结果的顺序性
  • 代码简洁,可读性高,维护成本低

总结

原实现的自旋等待和低效同步逻辑完全抵消了并行构建的优势,导致性能比串行还差。而利用Java并行流可以在保证输出顺序的前提下,最大化并行构建的效率,同时简化代码实现。

内容的提问来源于stack exchange,提问作者Anonymous Student

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 12:05:42