Selenium无法模拟Chrome的cache链接跳转行为求助
解决Selenium无法加载谷歌缓存页面的问题
问题根源
Chrome地址栏的cache:xxx是谷歌搜索的特殊指令,浏览器会自动将其解析为对应的缓存页面请求,但Selenium的driver.get()方法不识别这种非标准URL格式,只会把它当作无效地址处理,因此返回空页面。另外你原代码里遗漏了import time,运行时会触发NameError,这点也需要修正。
解决方法
方法1:直接使用谷歌缓存的标准URL
构造谷歌缓存的通用请求URL,替换掉代码中的cache:xxx格式地址:
from selenium import webdriver import time driver = webdriver.Chrome() try: # 构造通用的谷歌缓存请求URL cache_url = "http://webcache.googleusercontent.com/search?q=cache:https://www.nytimes.com/" driver.get(cache_url) time.sleep(5) print(driver.page_source) except Exception as X: print(X) driver.close()
这个URL不需要额外的参数(比如rlz、sourceid),谷歌会自动处理跳转。
方法2:模拟手动搜索缓存的流程
贴近Chrome手动操作的逻辑,先打开谷歌首页,再输入缓存指令并提交:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.common.keys import Keys import time driver = webdriver.Chrome() try: driver.get("https://www.google.com/") time.sleep(2) # 定位搜索框并输入缓存指令 search_box = driver.find_element(By.NAME, "q") search_box.send_keys("cache:https://www.nytimes.com/") search_box.send_keys(Keys.ENTER) time.sleep(5) print(driver.page_source) except Exception as X: print(X) driver.close()
额外注意事项
- 避免使用固定的
time.sleep(),推荐用WebDriverWait做显式等待,提升稳定性。 - 谷歌可能会检测到自动化工具,必要时可以给Chrome添加启动参数(如
--disable-blink-features=AutomationControlled)来绕过检测。
内容的提问来源于stack exchange,提问作者ruben
相关产品推荐
相关产品推荐

