求用Jsoup获取谷歌指定关键词全部搜索结果的示例代码
获取谷歌搜索结果(5000+条)的Jsoup实现示例
Hey there! Getting 5000+ Google search results with Jsoup is totally feasible, but first a critical heads-up: Google's Terms of Service strictly prohibit automated scraping of their search results at scale. If you proceed, you'll almost certainly run into anti-bot measures like CAPTCHAs, IP bans, or rate limits. So use this code only for educational purposes, and make sure you're fully compliant with Google's rules before moving forward.
That said, here's a sample implementation to help you understand the basics:
1. Add Jsoup Dependency
First, include Jsoup in your project. If you're using Maven, add this to your pom.xml:
<dependency> <groupId>org.jsoup</groupId> <artifactId>jsoup</artifactId> <version>1.17.2</version> <!-- Use the latest stable version --> </dependency>
2. Sample Java Code
import org.jsoup.Jsoup; import org.jsoup.nodes.Document; import org.jsoup.nodes.Element; import org.jsoup.select.Elements; import java.io.IOException; import java.util.HashSet; import java.util.Random; import java.util.Set; public class GoogleScraper { // Target keyword and desired number of results private static final String KEYWORD = "dentist las vegas"; private static final int TOTAL_RESULTS_NEEDED = 5000; // Google search base URL private static final String GOOGLE_SEARCH_URL = "https://www.google.com/search"; // Simulate a real browser user-agent to avoid immediate blocks private static final String USER_AGENT = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"; public static void main(String[] args) { Set<String> resultLinks = new HashSet<>(); Random random = new Random(); try { for (int start = 0; resultLinks.size() < TOTAL_RESULTS_NEEDED; start += 10) { // Construct the search URL with pagination (start parameter) String url = String.format("%s?q=%s&start=%d", GOOGLE_SEARCH_URL, KEYWORD.replace(" ", "+"), start); // Fetch the page with Jsoup, setting user-agent and a random delay Thread.sleep(random.nextInt(3000) + 2000); // 2-5 second delay between requests Document doc = Jsoup.connect(url) .userAgent(USER_AGENT) .timeout(10000) .get(); // Extract search result links - note: Google's HTML structure may change! Elements resultElements = doc.select("div.yuRUbf a"); // Current selector as of 2024 if (resultElements.isEmpty()) { System.out.println("No more results found or blocked by Google. Stopping..."); break; } // Add valid links to our set (avoid duplicates) for (Element link : resultElements) { String href = link.attr("href"); // Skip non-HTTP links (like Google's internal links) if (href.startsWith("http") && !href.contains("google.com")) { resultLinks.add(href); System.out.println("Added link: " + href); } } System.out.println(String.format("Collected %d/%d results so far...", resultLinks.size(), TOTAL_RESULTS_NEEDED)); } // Print final stats System.out.println(String.format("Final count: %d unique links collected.", resultLinks.size())); // You can write these links to a file or database here } catch (IOException e) { System.err.println("Error fetching page: " + e.getMessage()); } catch (InterruptedException e) { System.err.println("Thread interrupted: " + e.getMessage()); Thread.currentThread().interrupt(); } } }
3. Key Notes & Improvements
- Anti-Bot Mitigation: The code includes a random delay, but this won't be enough for 5000 results. You'll likely need:
- Proxy rotation (switch IPs regularly to avoid bans)
- A pool of random user-agent strings
- Avoiding rigid request patterns (don't hit the site every X seconds exactly)
- Selector Changes: Google frequently updates its HTML structure. If the selector
div.yuRUbf astops working, inspect the search page to find the new selector for result links. - Duplicate Prevention: Using a
HashSetensures we don't store duplicate links. - Compliance: Again, remember that large-scale scraping violates Google's Terms of Service. This is for educational use only.
内容的提问来源于stack exchange,提问作者Sonaldo
相关产品推荐
相关产品推荐

