Spring Boot中RequestMapping方法多线程优化Selenium爬取咨询
解决Spring Boot中Selenium爬取任务并行执行的问题
看起来你是因为串行执行Selenium爬取操作导致首页加载太慢,想要把这些耗时操作改成并行执行对吧?在Spring Boot里实现这个其实不难,但有几个关键要点需要注意,尤其是Selenium的线程安全问题,我给你两种可行的方案:
方案一:手动使用ExecutorService管理线程池
这种方式比较直接,适合快速改造现有代码:
步骤1:配置Spring线程池(推荐)
先在配置类里定义一个线程池Bean,统一管理线程资源,避免随意创建线程导致资源耗尽:
import org.springframework.context.annotation.Bean; import org.springframework.context.annotation.Configuration; import org.springframework.scheduling.concurrent.ThreadPoolTaskExecutor; import java.util.concurrent.ExecutorService; @Configuration public class ThreadPoolConfig { @Bean(name = "scrapeExecutor") public ExecutorService scrapeExecutor() { ThreadPoolTaskExecutor executor = new ThreadPoolTaskExecutor(); executor.setCorePoolSize(4); // 核心线程数,和你的爬取任务数对应 executor.setMaxPoolSize(8); // 最大线程数 executor.setQueueCapacity(10); executor.setThreadNamePrefix("scrape-thread-"); executor.initialize(); return executor.getThreadPoolExecutor(); } }
步骤2:改造Controller方法
在Controller里注入这个线程池,把爬取任务包装成Callable提交到线程池,用Future获取并行执行的结果:
@Autowired private WebScrape webscrape; @Autowired @Qualifier("scrapeExecutor") private ExecutorService scrapeExecutor; @RequestMapping(value = "/") public String printTable(ModelMap model) throws InterruptedException, ExecutionException { // 提交并行任务:把getAllData和getWorlValues拆成两个异步任务 Future<List<Object>> allDataFuture = scrapeExecutor.submit(() -> webscrape.getAllData()); Future<List<String>> worldValuesFuture = scrapeExecutor.submit(() -> webscrape.getWorlValues()); // 等待所有任务完成,获取结果 List<Object> allData = allDataFuture.get(); List<String> worldValues = worldValuesFuture.get(); // 组装页面数据 model.addAttribute("alldata", allData); model.addAttribute("worldCases", worldValues.get(0)); model.addAttribute("worldDeaths", worldValues.get(1)); model.addAttribute("worldPop", worldValues.get(2)); return "index"; }
如果getWorlValues()内部也是三个串行的爬取操作,你还可以进一步拆分它的内部逻辑,把三个子任务也改成并行,进一步提升效率。
方案二:使用Spring的@Async异步注解
这种方式更符合Spring的开发风格,代码更简洁:
步骤1:开启异步支持
在Spring Boot启动类上添加@EnableAsync注解,开启异步功能:
import org.springframework.boot.SpringApplication; import org.springframework.boot.autoconfigure.SpringBootApplication; import org.springframework.scheduling.annotation.EnableAsync; @SpringBootApplication @EnableAsync public class YourApplication { public static void main(String[] args) { SpringApplication.run(YourApplication.class, args); } }
步骤2:改造WebScrape类的方法
把需要并行执行的方法改成异步,返回CompletableFuture,重点:每个异步方法必须创建独立的WebDriver实例:
import org.springframework.scheduling.annotation.Async; import org.springframework.stereotype.Component; import java.util.concurrent.CompletableFuture; @Component public class WebScrape { // 每个任务创建独立的WebDriver,用完及时关闭 private WebDriver createWebDriver() { ChromeOptions options = new ChromeOptions(); options.addArguments("--headless"); // 无头模式大幅提升爬取效率 options.addArguments("--disable-gpu"); return new ChromeDriver(options); } @Async("scrapeExecutor") // 指定使用我们配置的线程池 public CompletableFuture<List<Object>> getAllData() { WebDriver driver = createWebDriver(); try { // 你的爬取逻辑,获取allData List<Object> data = new ArrayList<>(); // 填充数据逻辑... return CompletableFuture.completedFuture(data); } finally { driver.quit(); // 必须关闭driver释放资源 } } @Async("scrapeExecutor") public CompletableFuture<List<String>> getWorlValues() { WebDriver driver = createWebDriver(); try { // 你的爬取逻辑,获取cases、deaths、pop List<String> values = new ArrayList<>(); values.add("爬取的cases数据"); values.add("爬取的deaths数据"); values.add("爬取的pop数据"); return CompletableFuture.completedFuture(values); } finally { driver.quit(); } } }
这里一定要注意:Selenium的WebDriver是非线程安全的,多个线程绝对不能共享同一个实例,否则会出现元素找不到、页面跳转混乱等各种奇怪问题。
步骤3:改造Controller方法
现在可以异步调用这些方法,等待所有任务完成后组装结果:
@Autowired private WebScrape webscrape; @RequestMapping(value = "/") public String printTable(ModelMap model) { // 发起异步调用 CompletableFuture<List<Object>> allDataFuture = webscrape.getAllData(); CompletableFuture<List<String>> worldValuesFuture = webscrape.getWorlValues(); // 等待所有异步任务完成 CompletableFuture.allOf(allDataFuture, worldValuesFuture).join(); // 获取结果(这里可以加异常处理,避免任务失败导致页面报错) List<Object> allData = allDataFuture.join(); List<String> worldValues = worldValuesFuture.join(); // 组装页面数据 model.addAttribute("alldata", allData); model.addAttribute("worldCases", worldValues.get(0)); model.addAttribute("worldDeaths", worldValues.get(1)); model.addAttribute("worldPop", worldValues.get(2)); return "index"; }
必须注意的核心细节
- WebDriver资源管理:每个爬取任务必须使用独立的WebDriver实例,用完一定要调用
quit()关闭,避免浏览器进程残留占用资源。 - 异常处理:实际项目中要捕获异步任务的异常(比如爬取超时、页面加载失败),可以给用户返回默认数据或友好提示,避免页面直接崩溃。
- 无头模式:开启浏览器无头模式是提升爬取速度的关键,能减少大量资源占用,一定要加上。
- 线程池参数调整:根据服务器性能和爬取任务的耗时,调整线程池的核心线程数、最大线程数,避免线程过多导致服务器负载过高。
这样改造后,原本串行的爬取任务会并行执行,首页加载速度应该会有明显提升。
内容的提问来源于stack exchange,提问作者Fenzox
相关产品推荐
相关产品推荐

