Java并行有序字符串构建效率不及串行,矩阵打印遇性能瓶颈
并行打印矩阵到CLI的性能瓶颈解决
我有一个在while{}循环中持续更新的int[][]矩阵,每次迭代后需要将矩阵新值打印到CLI。为配合程序其他部分的并行逻辑,我实现了并行打印,但要求打印顺序严格对应矩阵行顺序。当前实现的性能比串行打印还慢,怀疑CountDownLatch的使用是瓶颈。
原实现代码
主循环代码
int numbOfBlocks = Runtime.getRuntime().availableProcessors(); int blockSize = matrix.length/numbOfBlocks; int endRow; while(true) { // print matrix to CLI StringBuilder stringBuilder = new StringBuilder(); CountDownLatch stringBuilderLatch = new CountDownLatch(numbOfBlocks); for (int i = 0; i < numbOfBlocks; i++) { endRow = (i == numbOfBlocks - 1) ? rows : (i+1) * blockSize; StringBuilderThread stringBuilderThread = new StringBuilderThread(matrix, i*blockSize, endRow, stringBuilderLatch, stringBuilder, i, numbOfBlocks); executor.execute(stringBuilderThread); } stringBuilderLatch.await(); stringBuilder.append("___"); System.out.println(stringBuilder.toString()); }
StringBuilderThread代码
import java.util.concurrent.CountDownLatch; public class StringBuilderThread implements Runnable { private int[][] matrixUpd; private int startRow; private int endRow; private CountDownLatch latch; private StringBuilder mainStringBuilder; private int order; private int numbOfBlocks; public StringBuilderThread(int[][] matrixUpd, int startRow, int endRow, CountDownLatch latch, StringBuilder mainStringBuilder, int order, int numbOfBlocks) { this.matrixUpd = matrixUpd; this.startRow = startRow; this.endRow = endRow; this.latch = latch; this.mainStringBuilder = mainStringBuilder; this.order = order; this.numbOfBlocks = numbOfBlocks; } @Override public void run() { StringBuilder tempStringBuilder = new StringBuilder(); int cols = matrixUpd[0].length; for (int i = startRow; i < endRow; i++) { tempStringBuilder.append("|"); for (int j = 0; j < cols; j++) { if(matrixUpd[i][j] == 1) { tempStringBuilder.append("* "); } else { tempStringBuilder.append(". "); } } tempStringBuilder.append("|\n"); } tempStringBuilder.append(" "); // order synchronization, so that threads write to stringbuilders in correct order long latchCount = latch.getCount(); while (!(numbOfBlocks-latchCount == order)) { latchCount = latch.getCount(); } mainStringBuilder.append(tempStringBuilder.toString()); latch.countDown(); } }
原实现的核心问题
原代码的性能瓶颈集中在自旋等待+共享StringBuilder的同步逻辑:
- 线程构建完子字符串后,通过持续轮询
CountDownLatch计数等待写入顺序,属于无意义的忙等待,占用大量CPU资源。 - 虽然线程用临时
StringBuilder构建局部内容,但最终写入共享mainStringBuilder时,自旋逻辑本质上强制了串行写入,完全抵消了并行构建的优势,还额外增加了线程调度和轮询的开销。
优化方案与测试结果
针对问题测试了三种解决方案,覆盖不同并行实现思路:
测试方案
- 串行打印:单线程遍历矩阵构建字符串并打印
- Callable并行构建:每个线程构建对应块的字符串,通过
Future收集结果后按顺序拼接 - 数组并行流构建:利用Java 8+的并行流,将矩阵行分块并行处理,再按顺序合并结果
测试场景
测试了四种不同维度的矩阵及对应迭代次数:
- 200x200,迭代100次
- 1000x1000,迭代100次
- 2000x2000,迭代200次
- 4000x4000,迭代200次
最优方案:数组并行流实现
并行流方案在所有测试场景中性能最优,且代码实现最简洁,无需手动管理线程和同步逻辑。示例代码如下:
while(true) { // 并行处理矩阵行,构建每行的字符串,再按顺序拼接 String matrixStr = IntStream.range(0, rows) .parallel() .mapToObj(i -> { StringBuilder rowBuilder = new StringBuilder("|"); for (int j = 0; j < cols; j++) { rowBuilder.append(matrix[i][j] == 1 ? "* " : ". "); } return rowBuilder.append("|\n").toString(); }) .collect(Collectors.joining()) + "___"; System.out.println(matrixStr); }
方案优势
- 并行流自动管理线程池(默认使用
ForkJoinPool.commonPool()),避免手动线程调度的开销 - 无需额外同步组件(如
CountDownLatch),通过流的collect操作自然保证结果的顺序性 - 代码简洁,可读性高,维护成本低
总结
原实现的自旋等待和低效同步逻辑完全抵消了并行构建的优势,导致性能比串行还差。而利用Java并行流可以在保证输出顺序的前提下,最大化并行构建的效率,同时简化代码实现。
内容的提问来源于stack exchange,提问作者Anonymous Student
相关产品推荐
相关产品推荐

