使用Jsoup网页爬取时,如何接受Google授权Cookie?
问题描述
我想用Jsoup库爬取谷歌上的英超积分榜页面,但访问时会遇到Google的Cookie授权页面,必须接受授权才能进入目标页面。网上没找到相关实现方法,我推测需要提交授权表单并存储Cookie,希望能得到帮助。
现有代码
import org.jsoup.Jsoup; import org.jsoup.nodes.Document; import org.jsoup.nodes.Element; import org.springframework.boot.SpringApplication; import org.springframework.boot.autoconfigure.SpringBootApplication; @SpringBootApplication public class ScoreTrackerApplication { public static void main(String[] args) { SpringApplication.run(ScoreTrackerApplication.class, args); final String url = "https://www.google.com/search?q=premier+league+fixtures&rlz=1C1CHBF_enIE874IE874&oq=premier+league+&aqs=chrome.2.69i57j0i131i433i512l3j0i433i512j0i131i433i512l5.8107j1j7&sourceid=chrome&ie=UTF-8#sie=lg;/g/11pz7zbpnb;2;/m/02_tc;st;fp;1;;;"; try { final Document document = Jsoup.connect(url).get(); //Testing purposes only System.out.println(document.body()); }catch (Exception ex){ } } }
爬取返回内容(中文翻译)
<h1 class="I90TVb" id="S3BnEe">继续访问Google前请完成验证</h1> <div class="AG96lb"> <div class="eLZYyf"> 我们使用<a class="F4a1l" href="https://policies.google.com/technologies/cookies?utm_source=ucbs&hl=en-IE" target="_blank">Cookie</a>和数据来: <ul class="dbXO9"> <li class="gowsYd ibCF0c">提供并维护Google服务</li> <li class="gowsYd GwwhGf">跟踪服务中断情况,防范垃圾信息、欺诈和滥用行为</li> <li class="gowsYd v8Bpfb">衡量受众参与度和网站统计数据,了解服务使用情况并提升服务质量</li> </ul> </div> <div class="eLZYyf"> 若您选择"全部接受",我们还会使用Cookie和数据来: <ul class="dbXO9"> <li class="gowsYd M6j9qf">开发和改进新服务</li> <li class="gowsYd v8Bpfb">投放并衡量广告效果</li> <li class="gowsYd e21Mac">根据您的设置展示个性化内容</li> <li class="gowsYd ohEWPc">根据您的设置展示个性化广告</li> </ul> <div class="jLhwdc"> 若您选择"全部拒绝",我们不会将Cookie用于上述额外用途。 </div> </div> <div class="yS1nld"> 非个性化内容受当前查看的内容、活跃搜索会话中的活动以及您所在位置等因素影响。非个性化广告受当前查看的内容和大致位置影响。个性化内容和广告还可能包含更相关的结果、推荐和定制广告,基于此浏览器的过往活动(如之前的Google搜索记录)。如有需要,我们还会使用Cookie和数据调整适合年龄段的体验。 </div> <div class="yS1nld"> 选择"更多选项"可查看额外信息,包括管理隐私设置的详情。您也可以随时访问<span>g.co/privacytools</span>。 </div> </div> </div> <div class="spoKVd"> <div class="GzLjMd"> <button id="W0wltc" class="tHlp8d" data-ved="0ahUKEwjVkZLx_939AhWObsAKHcCBBz4Q4cIICF8"> <div class="QS5gu sy4vM" role="none"> 全部拒绝 </div></button><button id="L2AGLb" class="tHlp8d" data-ved="0ahUKEwjVkZLx_939AhWObsAKHcCBBz4QiZAHCGA"> <div class="QS5gu sy4vM" role="none"> 全部接受 </div></button>
解决方案
要绕过Google的Cookie授权页面,核心是模拟用户点击"全部接受"按钮的请求,获取并保存授权Cookie,后续请求携带该Cookie访问目标页面。具体实现步骤如下:
1. 核心思路
Google的Cookie授权页面会生成包含授权参数的隐藏表单,点击"全部接受"时会提交POST请求并设置授权Cookie。我们需要:
- 先请求目标URL,获取授权页面和初始Cookie
- 提取授权所需的参数,构造POST请求提交授权
- 保存授权后的Cookie,携带它再次请求目标页面
2. 修改后的代码实现
import org.jsoup.Connection; import org.jsoup.Jsoup; import org.jsoup.nodes.Document; import org.springframework.boot.SpringApplication; import org.springframework.boot.autoconfigure.SpringBootApplication; import java.util.Map; @SpringBootApplication public class ScoreTrackerApplication { public static void main(String[] args) { SpringApplication.run(ScoreTrackerApplication.class, args); final String targetUrl = "https://www.google.com/search?q=premier+league+fixtures&rlz=1C1CHBF_enIE874IE874&oq=premier+league+&aqs=chrome.2.69i57j0i131i433i512l3j0i433i512j0i131i433i512l5.8107j1j7&sourceid=chrome&ie=UTF-8#sie=lg;/g/11pz7zbpnb;2;/m/02_tc;st;fp;1;;;"; Map<String, String> cookies = null; try { // 第一步:获取Cookie授权页面,保存初始Cookie Connection firstConn = Jsoup.connect(targetUrl) .userAgent("Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"); Document consentPage = firstConn.get(); cookies = firstConn.response().cookies(); // 提取授权所需参数(可从页面HTML中动态获取,此处为示例硬编码) String consentUrl = "https://consent.google.com/save"; String gl = "IE"; // 地区代码,根据页面内容调整 String hl = "en-IE"; // 语言代码,根据页面内容调整 String consent = "YES"; // 接受所有Cookie的标识 // 第二步:提交授权请求,更新Cookie Connection consentConn = Jsoup.connect(consentUrl) .userAgent("Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36") .cookies(cookies) .data("gl", gl) .data("hl", hl) .data("consent", consent) .data("continue", targetUrl) .method(Connection.Method.POST); consentConn.execute(); cookies = consentConn.response().cookies(); // 更新为包含授权的Cookie // 第三步:携带授权Cookie请求目标页面 Document targetDoc = Jsoup.connect(targetUrl) .userAgent("Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36") .cookies(cookies) .get(); // 输出目标页面内容(测试用) System.out.println(targetDoc.body()); } catch (Exception ex) { ex.printStackTrace(); } } }
3. 关键注意事项
- User-Agent必须设置:模拟真实浏览器的User-Agent,否则Google会识别为爬虫并拒绝服务
- 参数动态提取:
gl、hl等参数会随地区和语言变化,建议从授权页面的HTML中动态提取,避免硬编码失效 - 控制请求频率:Google对爬虫有严格的频率限制,短时间内多次请求会触发验证码或IP封禁
- 遵守服务条款:爬取Google内容需符合其官方服务条款,禁止用于商业或违规用途
内容的提问来源于stack exchange,提问作者Philip Herweling
相关产品推荐
相关产品推荐

