You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

JSoup新手技术问询:如何提取动态<a>标签及h2标签内的可变数据

Fixing Your JSoup Web Scraping for WhoIsMyTD.com

Hey there! I see you're new to JSoup and web scraping, and you're stuck trying to pull that constituency name (like Dublin Bay South) from WhoIsMyTD.com. Let's break down what's wrong with your current code and fix it step by step.

What's Wrong with Your Current Code?

  • Incorrect CSS Selector: Your selector well.col-md-4.h2 is invalid. well and col-md-4 are CSS classes (they need to start with a dot .), and h2 is a child element inside that container. The correct selector should be .well.col-md-4 h2.
  • Creating a Blank Element: You're making a new empty Element with new Element("well.col-md-4.h2") instead of using the element you actually scraped from the page. That's why you're getting an empty tag structure!
  • Using the Wrong Method to Get Text: toString() returns the full HTML of the element, not the text inside it. You need to use the text() method to extract the constituency name.

Fixed Code

private String jSoupTDRequest(String aLine1, String aLine3) throws IOException {
    String constit = "";
    String url = "https://www.whoismytd.com/search?utf8=✓&form-input="+aLine1+"%2C+"+aLine3+" +Ireland";
    
    Document doc = Jsoup.connect(url)
            .timeout(6000).get();
    
    // Correct selector to target the h2 inside the well.col-md-4 container
    Elements constituencyElements = doc.select(".well.col-md-4 h2");
    
    // Check if we found the element before accessing it (avoids null pointers)
    if (!constituencyElements.isEmpty()) {
        Element constituencyElement = constituencyElements.first();
        constit = constituencyElement.text(); // Get the text inside the h2 tag
    }
    
    return constit;
}

Bonus: Extracting from the Constituency URL

If you also need to get the constituency name directly from the URL path (like constituency/dublin-bay-south), you can do this since the search page redirects to the constituency page:

private String getConstituencyFromUrl(String aLine1, String aLine3) throws IOException {
    String url = "https://www.whoismytd.com/search?utf8=✓&form-input="+aLine1+"%2C+"+aLine3+" +Ireland";
    
    // Follow redirects (JSoup does this by default)
    Connection.Response response = Jsoup.connect(url)
            .timeout(6000).execute();
    
    String finalUrl = response.url().toString();
    String constituencyPath = finalUrl.split("constituency/")[1]; // Get part after "constituency/"
    
    // Convert kebab-case to title case (optional)
    String[] parts = constituencyPath.split("-");
    StringBuilder titleCase = new StringBuilder();
    for (String part : parts) {
        titleCase.append(Character.toUpperCase(part.charAt(0)))
                 .append(part.substring(1))
                 .append(" ");
    }
    
    return titleCase.toString().trim(); // Returns "Dublin Bay South"
}

Key Notes

  • Always check if your selected elements are empty before accessing them to avoid NullPointerException.
  • Use text() when you want plain text inside a tag, and html() if you need the inner HTML structure.
  • JSoup follows redirects by default, so you can easily grab the final URL after a search redirects to the constituency page.

内容的提问来源于stack exchange,提问作者BadHombe

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 06:37:30