使用Python+Selenium获取谷歌搜索结果href时遇NoSuchElementException求助
解决Selenium获取谷歌搜索结果href时的NoSuchElementException问题
错误原因分析
- 你用了
driver.find_element()(定位单个元素),却尝试用for循环遍历它,逻辑错误,应该用driver.find_elements()获取多个元素 - 谷歌搜索页面的DOM结构经常变动,固定层级的XPATH(比如
//*[@id="rso"]/div[1]/div/div/div[1]/div/a)极易失效 - 硬编码
time.sleep()等待页面加载不可靠,可能页面还没渲染完成就执行定位 - 未处理谷歌的Cookie同意弹窗,弹窗会遮挡或干扰元素定位
修正后的代码实现
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC text_fetch = "The Dutch Dress in Orange—Why?" url = f"https://www.google.com/search?q={text_fetch}" driver = webdriver.Chrome() driver.get(url) try: # 处理Cookie同意弹窗(如果存在) cookie_accept = WebDriverWait(driver, 10).until( EC.element_to_be_clickable((By.XPATH, '//div[text()="同意" or text()="Accept all"]')) ) cookie_accept.click() # 显式等待搜索结果加载完成,定位所有非广告的结果链接 Ggl_results = WebDriverWait(driver, 15).until( EC.presence_of_all_elements_located((By.CSS_SELECTOR, '#rso div.g a')) ) # 遍历获取每个结果的href for result in Ggl_results: href = result.get_attribute("href") if href and not href.startswith("https://www.google.com/aclk"): # 过滤广告链接 print(href) finally: driver.quit()
关键优化点
- 用
WebDriverWait显式等待替代time.sleep(),确保元素加载完成后再操作 - 改用更稳定的CSS选择器
#rso div.g a定位搜索结果链接,div.g是谷歌搜索结果的通用容器类 - 增加Cookie弹窗处理逻辑,避免弹窗干扰元素定位
- 过滤广告链接(广告链接通常以
https://www.google.com/aclk开头) - 用
find_elements()获取多个元素,支持遍历操作
内容的提问来源于stack exchange,提问作者Info Rewind
相关产品推荐
相关产品推荐

