Java将列表拆分为子列表后多线程并行处理的实现方案咨询
方案1:使用Java 8 Stream实现
你已经完成了列表拆分的步骤,直接对拆分后的子列表集合开启并行流处理即可,注意不要修改全局公共ForkJoinPool的参数,更推荐单独创建自定义线程池实现隔离:
import java.util.List; import java.util.concurrent.ForkJoinPool; import java.util.concurrent.ExecutionException; // 1. 拆分原始链接列表为固定大小的子列表 List<List<String>> partitionedLinks = Partition.ofSize(links, 3); // 2. 自定义并发度为5的ForkJoinPool,避免影响全局公共池 ForkJoinPool customPool = new ForkJoinPool(5); try { customPool.submit(() -> partitionedLinks.parallelStream().forEach(subList -> { // 每个子列表生成唯一的session编号,保证同子列表所有链接复用同一个代理会话 int sessionNum = subList.hashCode(); // 遍历当前子列表的所有链接,携带相同session调用爬取接口 subList.forEach(link -> ScrapeLinkWithSession(link, sessionNum)); }) ).get(); // 阻塞等待所有爬取任务执行完成 } catch (InterruptedException | ExecutionException e) { e.printStackTrace(); } finally { customPool.shutdown(); }
方案2:更优的CompletableFuture实现
如果你需要更灵活的任务控制(比如异常处理、超时控制、结果汇总),更推荐用CompletableFuture实现,扩展性更强:
import java.util.List; import java.util.concurrent.ExecutorService; import java.util.concurrent.Executors; import java.util.concurrent.CompletableFuture; import java.util.concurrent.atomic.AtomicInteger; import java.util.stream.Collectors; // 1. 拆分原始链接列表 List<List<String>> partitionedLinks = Partition.ofSize(links, 3); // 2. 创建固定大小的线程池,并发度设为5 ExecutorService executor = Executors.newFixedThreadPool(5); // 原子类生成全局唯一的session编号,避免hash冲突导致的会话复用问题 AtomicInteger sessionGenerator = new AtomicInteger(1000); // 3. 为每个子列表创建异步爬取任务 List<CompletableFuture<Void>> crawlTasks = partitionedLinks.stream() .map(subList -> CompletableFuture.runAsync(() -> { int currentSession = sessionGenerator.getAndIncrement(); subList.forEach(link -> ScrapeLinkWithSession(link, currentSession)); }, executor)) .collect(Collectors.toList()); // 4. 等待所有任务执行完成 CompletableFuture.allOf(crawlTasks.toArray(new CompletableFuture[0])).join(); executor.shutdown();
选择建议
- 如果你只需要基础的并行爬取,没有复杂的任务控制需求,选Stream方案即可,代码更简洁
- 如果你需要给爬取任务加超时、异常兜底、结果汇总等逻辑,选CompletableFuture方案,可控性更高
注意事项
- session编号要保证每个子列表唯一,避免不同子列表复用同一个会话导致代理逻辑混乱
- 线程池并发度需要根据ScraperAPI的限流规则调整,不要超过服务商的并发限制
- 如果爬取过程IO等待时间较长,可以适当调大线程池大小,提升资源利用率
内容的提问来源于stack exchange,提问作者user15676007
相关产品推荐
相关产品推荐

