You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Jsoup解析谷歌天气页面遇ERROR 429问题求助

解决Jsoup请求Google天气触发ERROR 429的问题

问题描述

用Java结合Jsoup开发简易天气解析器,昨日运行正常,今日立即出现ERROR 429(推测是请求次数过多触发反爬机制)。已尝试在Jsoup.connect()中添加.timeout()方法,但无效果。代码如下:

public class Parser {
    public static void parse() throws IOException {
        Scanner scanner = new Scanner(System.in);
        
        String city, temp, hum, wind, status, name, day;
        city = scanner.nextLine();

        Document doc = Jsoup.connect("https://www.google.com/search?q="+city+" weather").timeout(10*1000).get();
        Element tempElem = doc.selectFirst("span.wob_t.q8U8x");

        temp = Objects.requireNonNull(doc.selectFirst("span.wob_t.q8U8x")).text();
        hum = Objects.requireNonNull(doc.selectFirst("#wob_hm")).text();
        wind = Objects.requireNonNull(doc.selectFirst("#wob_ws")).text();
        status = Objects.requireNonNull(doc.selectFirst("#wob_dc")).text();
        name = Objects.requireNonNull(doc.selectFirst("#wob_loc.q8U8x")).text();
        day = Objects.requireNonNull(doc.selectFirst("#wob_dts")).text();
        
        if(tempElem == null){
            System.out.println("City's not found");
            System.exit(0);
        }

        System.out.println("Weather in " + name + ". ("+day+", "+status+")"+
                "\n Temperature: " + temp + "°C" +
                "\n Humanity: " + hum +
                "\n Wind speed:  " + wind);
    }
}

解决方案

ERROR 429是服务器返回的"请求过多"响应,和超时无关,核心是要规避Google的反爬检测,同时优化代码逻辑:

1. 添加浏览器请求头,模拟正常用户访问

Google会识别非浏览器发起的请求,给Jsoup请求添加User-Agent等头信息,伪装成浏览器:

Document doc = Jsoup.connect("https://www.google.com/search?q="+city+" weather")
    // 模拟Chrome浏览器请求头,可根据自己浏览器实际UA替换
    .userAgent("Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36")
    .timeout(10*1000)
    .get();

2. 增加请求间隔,降低请求频率

如果是短时间内频繁请求导致的限制,每次请求完成后添加随机延迟:

// 在parse方法末尾添加延迟,随机1-3秒
try {
    Thread.sleep((long)(Math.random() * 2000) + 1000);
} catch (InterruptedException e) {
    // 恢复中断状态
    Thread.currentThread().interrupt();
}

3. 缓存查询结果,避免重复请求

对已经查询过的城市天气结果进行缓存(比如用HashMap存储),短时间内重复查询直接返回缓存内容,减少请求次数:

// 类级别缓存
private static final Map<String, String> WEATHER_CACHE = new HashMap<>();

public static void parse() throws IOException {
    Scanner scanner = new Scanner(System.in);
    String city = scanner.nextLine();

    // 先查缓存
    if(WEATHER_CACHE.containsKey(city)){
        System.out.println(WEATHER_CACHE.get(city));
        return;
    }

    // 下面是原请求逻辑...

    // 组装结果并存入缓存
    String result = "Weather in " + name + ". ("+day+", "+status+")"+
            "\n Temperature: " + temp + "°C" +
            "\n Humanity: " + hum +
            "\n Wind speed:  " + wind;
    WEATHER_CACHE.put(city, result);
    System.out.println(result);
}

4. 修复代码逻辑漏洞

原代码中tempElem的null判断在元素获取之后,会导致Objects.requireNonNull先抛出空指针异常,需要调整顺序;同时复用已获取的元素,减少DOM查询:

Document doc = Jsoup.connect("https://www.google.com/search?q="+city+" weather")
        .userAgent("Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36")
        .timeout(10*1000)
        .get();
Element tempElem = doc.selectFirst("span.wob_t.q8U8x");

// 先判断城市是否存在
if(tempElem == null){
    System.out.println("City's not found");
    System.exit(0);
}

// 复用已获取的tempElem,避免重复查询DOM
temp = tempElem.text();
hum = Objects.requireNonNull(doc.selectFirst("#wob_hm")).text();
wind = Objects.requireNonNull(doc.selectFirst("#wob_ws")).text();
status = Objects.requireNonNull(doc.selectFirst("#wob_dc")).text();
name = Objects.requireNonNull(doc.selectFirst("#wob_loc.q8U8x")).text();
day = Objects.requireNonNull(doc.selectFirst("#wob_dts")).text();

内容的提问来源于stack exchange,提问作者Ohonovskiy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.07 05:01:41