使用Java ReadWriteLock实现多线程网页爬虫:如何避免重复爬取URL
解决多线程网页爬虫重复爬取URL的问题
你的代码出现重复爬取的核心原因是检查URL是否已访问与将URL标记为已访问的操作不是原子的,多个线程可能同时通过contains()的检查,进而重复爬取同一个URL。另外代码中还存在多实例共享的问题,下面是具体的修复方案:
1. 原子化URL的检查与添加操作
把contains()和addUrlToSeenURLSet()合并成一个原子方法,确保只有第一个线程能成功标记URL为已访问,后续线程直接跳过。利用ReadWriteLock的写锁保证原子性:
// 替换原来的contains和addUrlToSeenURLSet方法 public boolean markAsSeen(String url) { writeLock_http_request.lock(); try { // 利用HashSet.add()的返回值:添加成功返回true,已存在返回false return seenURL.add(url); } finally { writeLock_http_request.unlock(); } }
这个方法将检查和添加合并为一个原子操作,彻底避免了两步操作之间的线程切换导致的重复问题。
2. 修改crawl方法逻辑
使用新的原子方法替换原来的两步操作:
public void crawl(String startUrl, HtmlParser htmlParser) throws Exception { // 只有成功标记为已访问时,才执行爬取 if (markAsSeen(startUrl)) { try { List<String> subUrls = htmlParser.getUrls(startUrl); resultVisitedUrls.add(startUrl + " Done by thread - " + Thread.currentThread()); // 若要大规模爬取,建议将子URL任务提交到线程池,而非递归调用 for (String subUrl : subUrls) { crawl(subUrl, htmlParser); } } catch (Exception ex) { throw new Exception("Something went wrong. Method - crawl : " + ex.getMessage()); } } }
3. 统一使用单个WebCrawler实例
你的main方法中创建了多个WebCrawler实例,虽然seenURL是static的,但resultVisitedUrls是实例变量,会导致结果分散。所有任务应共享同一个实例:
public static void main(String[] args) { class Crawl implements Callable<List<String>> { String startUrl; WebCrawler webCrawler; public Crawl(String startUrl, WebCrawler webCrawler){ this.startUrl = startUrl; this.webCrawler = webCrawler; } public List<String> call() { HtmlParser htmlParser = new RetrieveURLs(); try { webCrawler.crawl(startUrl, htmlParser); } catch(Exception ex){ System.err.println("Some exception occurred in call() - " + ex.getMessage()); } return webCrawler.getResultVisitedUrls(); } } ExecutorService service = Executors.newFixedThreadPool(4); try{ // 所有任务共享同一个爬虫实例 WebCrawler sharedCrawler = new WebCrawler(); Future<List<String>> future_1 = service.submit(new Crawl("http://localhost:3001/getUrls/google.com", sharedCrawler)); Future<List<String>> future_2 = service.submit(new Crawl("http://localhost:3001/getUrls/google.com", sharedCrawler)); Future<List<String>> future_3 = service.submit(new Crawl("http://localhost:3001/getUrls/google.com", sharedCrawler)); // 共享实例的结果已汇总,无需合并多个返回结果 List<String> finalResult = sharedCrawler.getResultVisitedUrls(); System.out.println(finalResult.size()); for(String str : finalResult){ System.out.println(str ); } } catch(ExecutionException | InterruptedException ex){ ex.printStackTrace(); } finally { service.shutdown(); } }
4. 额外优化建议
- 大规模爬取时,递归调用
crawl易导致线程栈溢出,建议将子URL的爬取任务提交到ExecutorService执行。 - 可以使用
ConcurrentHashMap.newKeySet()替代HashSet加锁,它本身就是线程安全的,能简化代码:
这样可以完全省去ReadWriteLock的使用,利用JUC提供的线程安全集合保证原子性。static Set<String> seenURL = ConcurrentHashMap.newKeySet(); // 此时markAsSeen方法可简化为: public boolean markAsSeen(String url) { return seenURL.add(url); }
内容的提问来源于stack exchange,提问作者Karthik Kolathumani
相关产品推荐
相关产品推荐

