如何将Scrapy Splash获取的谷歌登录Cookie传递给Selenium绕过登录
修复方案
核心问题说明
Splash同步Cookie到ChromeDriver后仍跳登录,主要是域处理逻辑错误、Cookie字段不匹配、ChromeDriver自动化特征被谷歌识别三个原因导致,以下是可直接运行的修复方案:
1. 修复Cookie域处理逻辑
你的现有代码会把前缀带.的域自动拼接www,导致Cookie被写入错误域名下,谷歌的跨域Cookie需要写入对应根域,不需要额外加前缀,正确逻辑是先访问一次谷歌根域再批量添加Cookie。
2. 适配Splash与Selenium的Cookie字段
Splash返回的Cookie字段和Selenium要求的标准字段存在差异,需要先做字段过滤转换,避免无效字段导致Cookie添加失败。
3. 屏蔽ChromeDriver自动化特征
谷歌会识别默认启动的ChromeDriver的自动化标识,哪怕Cookie正确也会强制触发登录校验,需要添加启动参数隐藏自动化特征。
4. 完整修改后的parse方法代码
def parse(self, response): imgdata = base64.b64decode(response.data['png']) with open('image.png', 'wb') as file: file.write(imgdata) cookies = response.data.get("cookies") # 初始化带反检测配置的ChromeDriver options = webdriver.ChromeOptions() options.add_argument("--disable-blink-features=AutomationControlled") options.add_experimental_option("excludeSwitches", ["enable-automation"]) options.add_experimental_option("useAutomationExtension", False) driver = webdriver.Chrome("./chromedriver", options=options) # 修改window.navigator.webdriver属性绕过检测 driver.execute_cdp_cmd("Page.addScriptToEvaluateOnNewDocument", { "source": "Object.defineProperty(navigator, 'webdriver', {get: () => undefined})" }) # 先访问谷歌账号根域,为添加Cookie做准备 driver.get("https://accounts.google.com") time.sleep(2) # 转换并添加Cookie for cookie in cookies: # 适配Selenium要求的Cookie字段 formatted_cookie = { "name": cookie["name"], "value": cookie["value"], "domain": cookie["domain"], "path": cookie.get("path", "/"), "secure": cookie.get("secure", True), } # 过期时间存在才添加,避免格式报错 if "expiry" in cookie: formatted_cookie["expiry"] = cookie["expiry"] try: driver.add_cookie(formatted_cookie) except Exception as e: # 忽略无效域的Cookie报错 pass # 刷新页面验证登录状态 driver.get("https://accounts.google.com") time.sleep(3) # 再跳转目标控制台页面 driver.get("https://console.cloud.google.com/apis/library/youtube.googleapis.com") time.sleep(5)
内容的提问来源于stack exchange,提问作者Ubaid
相关产品推荐
相关产品推荐

