You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

WebCrawler爬虫仅抓取部分.html链接,请求排查问题根源

爬虫无法捕获Yahoo页面全部.html链接的问题排查

我编写了以下WebCrawler爬虫类:

public class WebCrawler {

    private HashSet<String> visitedUrls;

    public WebCrawler() {
        visitedUrls = new HashSet<String>();
    }

    public void crawl(String url) throws IOException {
        // Check if URL is already visited
        Document doc = Jsoup.connect(url).get();
        Set<String> links = new HashSet<String>();
        URL urlObject = new URL(url);
        String domainName = urlObject.getHost();
        Elements elements = doc.select("a[href]");
        for (Element e : elements) {
            String l = e.attr("href");
            if(l.endsWith(".html")){
                if(!l.startsWith("http")){
                    l = domainName + l;
                    // Determine whether the URL is HTTP or HTTPS
                    if (urlObject.getProtocol().equals("https")) {
                       l = "https:/" + l;
                    } else if (urlObject.getProtocol().equals("http")) {
                            l = "http:/" + l;
                    } else {
                        System.out.println("The URL is neither HTTP nor HTTPS.");
                    }
                    links.add(l);
                    System.out.println(l);
                }

            }

        }
    }

    public static void main(String[] args) throws IOException {
        WebCrawler crawler = new WebCrawler();
        String url = "...";  // <- replace a URL
        crawler.crawl(url);
    }
}

我在Yahoo的PRVB新闻页面测试该代码时,仅能检索到前四个以.html结尾的链接,无法捕获页面上的全部链接,请问问题出在哪里?


问题根源分析

  1. 动态内容未加载:Yahoo新闻页面的大量链接是通过JavaScript动态渲染的,Jsoup只能获取页面初始加载的HTML源码,无法执行JS,因此滚动或交互后才出现的链接无法被抓取到。
  2. URL拼接错误:代码中拼接相对路径时使用了https:/或http:/(缺少一个斜杠),正确格式应为https://或http://,这会导致生成的URL无效,部分符合条件的链接被错误处理后未加入集合。
  3. 过滤条件过于严苛:仅筛选.html结尾的链接,但部分新闻链接可能带有查询参数(比如?param=xxx),或是重定向到.html页面的非.html链接,这些都会被当前代码过滤掉。
  4. 未启用已访问URL校验:虽然定义了visitedUrls集合,但代码中并未在爬取前检查URL是否已访问,不过这并非首次爬取就拿不全链接的直接原因。

解决建议

  • 处理动态内容:改用Selenium、Playwright等支持JavaScript执行的工具,模拟浏览器完整加载页面后再抓取链接。
  • 修复URL拼接:将代码中的https:/改为https://,http:/改为http://,确保生成的URL格式正确。
  • 放宽过滤规则:先收集所有<a>标签的链接,再通过判断域名、路径特征等方式筛选目标链接,而非仅依赖.html后缀。
  • 启用已访问校验:在crawl方法开头添加if (visitedUrls.contains(url)) return;,并在获取页面后将URL加入visitedUrls,避免重复爬取。

内容的提问来源于stack exchange,提问作者vic

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 04:23:14