使用Jsoup解析谷歌天气页面遇ERROR 429问题求助
解决Jsoup请求Google天气触发ERROR 429的问题
问题描述
用Java结合Jsoup开发简易天气解析器,昨日运行正常,今日立即出现ERROR 429(推测是请求次数过多触发反爬机制)。已尝试在Jsoup.connect()中添加.timeout()方法,但无效果。代码如下:
public class Parser { public static void parse() throws IOException { Scanner scanner = new Scanner(System.in); String city, temp, hum, wind, status, name, day; city = scanner.nextLine(); Document doc = Jsoup.connect("https://www.google.com/search?q="+city+" weather").timeout(10*1000).get(); Element tempElem = doc.selectFirst("span.wob_t.q8U8x"); temp = Objects.requireNonNull(doc.selectFirst("span.wob_t.q8U8x")).text(); hum = Objects.requireNonNull(doc.selectFirst("#wob_hm")).text(); wind = Objects.requireNonNull(doc.selectFirst("#wob_ws")).text(); status = Objects.requireNonNull(doc.selectFirst("#wob_dc")).text(); name = Objects.requireNonNull(doc.selectFirst("#wob_loc.q8U8x")).text(); day = Objects.requireNonNull(doc.selectFirst("#wob_dts")).text(); if(tempElem == null){ System.out.println("City's not found"); System.exit(0); } System.out.println("Weather in " + name + ". ("+day+", "+status+")"+ "\n Temperature: " + temp + "°C" + "\n Humanity: " + hum + "\n Wind speed: " + wind); } }
解决方案
ERROR 429是服务器返回的"请求过多"响应,和超时无关,核心是要规避Google的反爬检测,同时优化代码逻辑:
1. 添加浏览器请求头,模拟正常用户访问
Google会识别非浏览器发起的请求,给Jsoup请求添加User-Agent等头信息,伪装成浏览器:
Document doc = Jsoup.connect("https://www.google.com/search?q="+city+" weather") // 模拟Chrome浏览器请求头,可根据自己浏览器实际UA替换 .userAgent("Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36") .timeout(10*1000) .get();
2. 增加请求间隔,降低请求频率
如果是短时间内频繁请求导致的限制,每次请求完成后添加随机延迟:
// 在parse方法末尾添加延迟,随机1-3秒 try { Thread.sleep((long)(Math.random() * 2000) + 1000); } catch (InterruptedException e) { // 恢复中断状态 Thread.currentThread().interrupt(); }
3. 缓存查询结果,避免重复请求
对已经查询过的城市天气结果进行缓存(比如用HashMap存储),短时间内重复查询直接返回缓存内容,减少请求次数:
// 类级别缓存 private static final Map<String, String> WEATHER_CACHE = new HashMap<>(); public static void parse() throws IOException { Scanner scanner = new Scanner(System.in); String city = scanner.nextLine(); // 先查缓存 if(WEATHER_CACHE.containsKey(city)){ System.out.println(WEATHER_CACHE.get(city)); return; } // 下面是原请求逻辑... // 组装结果并存入缓存 String result = "Weather in " + name + ". ("+day+", "+status+")"+ "\n Temperature: " + temp + "°C" + "\n Humanity: " + hum + "\n Wind speed: " + wind; WEATHER_CACHE.put(city, result); System.out.println(result); }
4. 修复代码逻辑漏洞
原代码中tempElem的null判断在元素获取之后,会导致Objects.requireNonNull先抛出空指针异常,需要调整顺序;同时复用已获取的元素,减少DOM查询:
Document doc = Jsoup.connect("https://www.google.com/search?q="+city+" weather") .userAgent("Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36") .timeout(10*1000) .get(); Element tempElem = doc.selectFirst("span.wob_t.q8U8x"); // 先判断城市是否存在 if(tempElem == null){ System.out.println("City's not found"); System.exit(0); } // 复用已获取的tempElem,避免重复查询DOM temp = tempElem.text(); hum = Objects.requireNonNull(doc.selectFirst("#wob_hm")).text(); wind = Objects.requireNonNull(doc.selectFirst("#wob_ws")).text(); status = Objects.requireNonNull(doc.selectFirst("#wob_dc")).text(); name = Objects.requireNonNull(doc.selectFirst("#wob_loc.q8U8x")).text(); day = Objects.requireNonNull(doc.selectFirst("#wob_dts")).text();
内容的提问来源于stack exchange,提问作者Ohonovskiy
相关产品推荐
相关产品推荐

