You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Java将列表拆分为子列表后多线程并行处理的实现方案咨询

方案1:使用Java 8 Stream实现

你已经完成了列表拆分的步骤,直接对拆分后的子列表集合开启并行流处理即可,注意不要修改全局公共ForkJoinPool的参数,更推荐单独创建自定义线程池实现隔离:

import java.util.List;
import java.util.concurrent.ForkJoinPool;
import java.util.concurrent.ExecutionException;

// 1. 拆分原始链接列表为固定大小的子列表
List<List<String>> partitionedLinks = Partition.ofSize(links, 3);
// 2. 自定义并发度为5的ForkJoinPool,避免影响全局公共池
ForkJoinPool customPool = new ForkJoinPool(5);

try {
    customPool.submit(() ->
        partitionedLinks.parallelStream().forEach(subList -> {
            // 每个子列表生成唯一的session编号,保证同子列表所有链接复用同一个代理会话
            int sessionNum = subList.hashCode();
            // 遍历当前子列表的所有链接,携带相同session调用爬取接口
            subList.forEach(link -> ScrapeLinkWithSession(link, sessionNum));
        })
    ).get(); // 阻塞等待所有爬取任务执行完成
} catch (InterruptedException | ExecutionException e) {
    e.printStackTrace();
} finally {
    customPool.shutdown();
}

方案2:更优的CompletableFuture实现

如果你需要更灵活的任务控制(比如异常处理、超时控制、结果汇总),更推荐用CompletableFuture实现,扩展性更强:

import java.util.List;
import java.util.concurrent.ExecutorService;
import java.util.concurrent.Executors;
import java.util.concurrent.CompletableFuture;
import java.util.concurrent.atomic.AtomicInteger;
import java.util.stream.Collectors;

// 1. 拆分原始链接列表
List<List<String>> partitionedLinks = Partition.ofSize(links, 3);
// 2. 创建固定大小的线程池,并发度设为5
ExecutorService executor = Executors.newFixedThreadPool(5);
// 原子类生成全局唯一的session编号,避免hash冲突导致的会话复用问题
AtomicInteger sessionGenerator = new AtomicInteger(1000);

// 3. 为每个子列表创建异步爬取任务
List<CompletableFuture<Void>> crawlTasks = partitionedLinks.stream()
        .map(subList -> CompletableFuture.runAsync(() -> {
            int currentSession = sessionGenerator.getAndIncrement();
            subList.forEach(link -> ScrapeLinkWithSession(link, currentSession));
        }, executor))
        .collect(Collectors.toList());

// 4. 等待所有任务执行完成
CompletableFuture.allOf(crawlTasks.toArray(new CompletableFuture[0])).join();
executor.shutdown();

选择建议

  • 如果你只需要基础的并行爬取,没有复杂的任务控制需求,选Stream方案即可,代码更简洁
  • 如果你需要给爬取任务加超时、异常兜底、结果汇总等逻辑,选CompletableFuture方案,可控性更高

注意事项

  • session编号要保证每个子列表唯一,避免不同子列表复用同一个会话导致代理逻辑混乱
  • 线程池并发度需要根据ScraperAPI的限流规则调整,不要超过服务商的并发限制
  • 如果爬取过程IO等待时间较长,可以适当调大线程池大小,提升资源利用率

内容的提问来源于stack exchange,提问作者user15676007

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.29 21:39:00