JSoup新手技术问询:如何提取动态<a>标签及h2标签内的可变数据
Fixing Your JSoup Web Scraping for WhoIsMyTD.com
Hey there! I see you're new to JSoup and web scraping, and you're stuck trying to pull that constituency name (like Dublin Bay South) from WhoIsMyTD.com. Let's break down what's wrong with your current code and fix it step by step.
What's Wrong with Your Current Code?
- Incorrect CSS Selector: Your selector
well.col-md-4.h2is invalid.wellandcol-md-4are CSS classes (they need to start with a dot.), andh2is a child element inside that container. The correct selector should be.well.col-md-4 h2. - Creating a Blank Element: You're making a new empty
Elementwithnew Element("well.col-md-4.h2")instead of using the element you actually scraped from the page. That's why you're getting an empty tag structure! - Using the Wrong Method to Get Text:
toString()returns the full HTML of the element, not the text inside it. You need to use thetext()method to extract the constituency name.
Fixed Code
private String jSoupTDRequest(String aLine1, String aLine3) throws IOException { String constit = ""; String url = "https://www.whoismytd.com/search?utf8=✓&form-input="+aLine1+"%2C+"+aLine3+" +Ireland"; Document doc = Jsoup.connect(url) .timeout(6000).get(); // Correct selector to target the h2 inside the well.col-md-4 container Elements constituencyElements = doc.select(".well.col-md-4 h2"); // Check if we found the element before accessing it (avoids null pointers) if (!constituencyElements.isEmpty()) { Element constituencyElement = constituencyElements.first(); constit = constituencyElement.text(); // Get the text inside the h2 tag } return constit; }
Bonus: Extracting from the Constituency URL
If you also need to get the constituency name directly from the URL path (like constituency/dublin-bay-south), you can do this since the search page redirects to the constituency page:
private String getConstituencyFromUrl(String aLine1, String aLine3) throws IOException { String url = "https://www.whoismytd.com/search?utf8=✓&form-input="+aLine1+"%2C+"+aLine3+" +Ireland"; // Follow redirects (JSoup does this by default) Connection.Response response = Jsoup.connect(url) .timeout(6000).execute(); String finalUrl = response.url().toString(); String constituencyPath = finalUrl.split("constituency/")[1]; // Get part after "constituency/" // Convert kebab-case to title case (optional) String[] parts = constituencyPath.split("-"); StringBuilder titleCase = new StringBuilder(); for (String part : parts) { titleCase.append(Character.toUpperCase(part.charAt(0))) .append(part.substring(1)) .append(" "); } return titleCase.toString().trim(); // Returns "Dublin Bay South" }
Key Notes
- Always check if your selected elements are empty before accessing them to avoid
NullPointerException. - Use
text()when you want plain text inside a tag, andhtml()if you need the inner HTML structure. - JSoup follows redirects by default, so you can easily grab the final URL after a search redirects to the constituency page.
内容的提问来源于stack exchange,提问作者BadHombe
相关产品推荐
相关产品推荐

