You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求用Jsoup获取谷歌指定关键词全部搜索结果的示例代码

获取谷歌搜索结果(5000+条)的Jsoup实现示例

Hey there! Getting 5000+ Google search results with Jsoup is totally feasible, but first a critical heads-up: Google's Terms of Service strictly prohibit automated scraping of their search results at scale. If you proceed, you'll almost certainly run into anti-bot measures like CAPTCHAs, IP bans, or rate limits. So use this code only for educational purposes, and make sure you're fully compliant with Google's rules before moving forward.

That said, here's a sample implementation to help you understand the basics:

1. Add Jsoup Dependency

First, include Jsoup in your project. If you're using Maven, add this to your pom.xml:

<dependency>
    <groupId>org.jsoup</groupId>
    <artifactId>jsoup</artifactId>
    <version>1.17.2</version> <!-- Use the latest stable version -->
</dependency>

2. Sample Java Code

import org.jsoup.Jsoup;
import org.jsoup.nodes.Document;
import org.jsoup.nodes.Element;
import org.jsoup.select.Elements;

import java.io.IOException;
import java.util.HashSet;
import java.util.Random;
import java.util.Set;

public class GoogleScraper {
    // Target keyword and desired number of results
    private static final String KEYWORD = "dentist las vegas";
    private static final int TOTAL_RESULTS_NEEDED = 5000;
    // Google search base URL
    private static final String GOOGLE_SEARCH_URL = "https://www.google.com/search";
    // Simulate a real browser user-agent to avoid immediate blocks
    private static final String USER_AGENT = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36";

    public static void main(String[] args) {
        Set<String> resultLinks = new HashSet<>();
        Random random = new Random();

        try {
            for (int start = 0; resultLinks.size() < TOTAL_RESULTS_NEEDED; start += 10) {
                // Construct the search URL with pagination (start parameter)
                String url = String.format("%s?q=%s&start=%d", GOOGLE_SEARCH_URL, KEYWORD.replace(" ", "+"), start);

                // Fetch the page with Jsoup, setting user-agent and a random delay
                Thread.sleep(random.nextInt(3000) + 2000); // 2-5 second delay between requests
                Document doc = Jsoup.connect(url)
                        .userAgent(USER_AGENT)
                        .timeout(10000)
                        .get();

                // Extract search result links - note: Google's HTML structure may change!
                Elements resultElements = doc.select("div.yuRUbf a"); // Current selector as of 2024

                if (resultElements.isEmpty()) {
                    System.out.println("No more results found or blocked by Google. Stopping...");
                    break;
                }

                // Add valid links to our set (avoid duplicates)
                for (Element link : resultElements) {
                    String href = link.attr("href");
                    // Skip non-HTTP links (like Google's internal links)
                    if (href.startsWith("http") && !href.contains("google.com")) {
                        resultLinks.add(href);
                        System.out.println("Added link: " + href);
                    }
                }

                System.out.println(String.format("Collected %d/%d results so far...", resultLinks.size(), TOTAL_RESULTS_NEEDED));
            }

            // Print final stats
            System.out.println(String.format("Final count: %d unique links collected.", resultLinks.size()));
            // You can write these links to a file or database here
        } catch (IOException e) {
            System.err.println("Error fetching page: " + e.getMessage());
        } catch (InterruptedException e) {
            System.err.println("Thread interrupted: " + e.getMessage());
            Thread.currentThread().interrupt();
        }
    }
}

3. Key Notes & Improvements

  • Anti-Bot Mitigation: The code includes a random delay, but this won't be enough for 5000 results. You'll likely need:
    • Proxy rotation (switch IPs regularly to avoid bans)
    • A pool of random user-agent strings
    • Avoiding rigid request patterns (don't hit the site every X seconds exactly)
  • Selector Changes: Google frequently updates its HTML structure. If the selector div.yuRUbf a stops working, inspect the search page to find the new selector for result links.
  • Duplicate Prevention: Using a HashSet ensures we don't store duplicate links.
  • Compliance: Again, remember that large-scale scraping violates Google's Terms of Service. This is for educational use only.

内容的提问来源于stack exchange,提问作者Sonaldo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:16:36