使用BeautifulSoup、Selenium爬取bmstores站点详情按钮链接报错求助
问题解决方案
报错核心原因
- 代码缺少目标页面加载逻辑,浏览器启动后没有访问
https://www.bmstores.co.uk/stores?location=KA8+9BF,直接查找元素必然触发NoSuchElementException - 绝对XPath容错性极低,页面DOM结构稍有变动就会定位失败
- 未提前处理站点的Cookie同意弹窗,弹窗遮挡目标元素会触发
ElementNotInteractableException - 硬等待
time.sleep不稳定,容易出现元素还未加载完成就执行查找操作的问题 - 你用到的
find_element_by_xpath、switch_to_window都是Selenium 3的废弃方法,4.0+版本已移除这类接口,需要用新的定位方式
修正后可运行代码
from selenium import webdriver as wd from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.common.exceptions import TimeoutException, NoSuchElementException local_path_of_chrome_driver = "E:\\chromedriver.exe" # 适配Selenium 4+的启动配置 options = wd.ChromeOptions() options.add_argument("--start-maximized") # 可选:禁用弹窗减少干扰 options.add_argument("--disable-notifications") driver = wd.Chrome(executable_path=local_path_of_chrome_driver, options=options) wait = WebDriverWait(driver, 10) data_links=[] try: # 第一步:加载目标页面 driver.get("https://www.bmstores.co.uk/stores?location=KA8+9BF") # 第二步:处理Cookie同意弹窗,避免遮挡元素 try: accept_cookie_btn = wait.until(EC.element_to_be_clickable((By.ID, "onetrust-accept-btn-handler"))) accept_cookie_btn.click() except (TimeoutException, NoSuchElementException): # 没找到弹窗就跳过,不影响后续逻辑 pass # 第三步:定位所有View Details按钮,用相对xpath匹配文本,稳定性更高 view_details_btns = wait.until(EC.presence_of_all_elements_located((By.XPATH, "//a[normalize-space(text())='View details']"))) # 直接提取href属性即可,不需要跳转页面,效率更高 for btn in view_details_btns: link = btn.get_attribute("href") if link and link not in data_links: data_links.append(link) print("爬取到的详情链接列表:") for link in data_links: print(link) finally: driver.quit()
代码优化点说明
- 替换绝对XPath为匹配文本的相对XPath,只要按钮文本不变,页面结构调整也不会影响定位
- 用显式等待替代硬等待,最多等10秒,元素加载完成就立即执行后续操作,兼顾稳定性和效率
- 提前处理Cookie弹窗,解决元素被遮挡无法交互的问题
- 直接提取
href属性获取链接,不需要切换窗口、后退等操作,大幅降低逻辑复杂度和报错概率
内容的提问来源于stack exchange,提问作者Jagadeesan jack
相关产品推荐
相关产品推荐

