使用Selenium网页爬取时点击元素报错的问题求助
解决NOAA气候数据爬取中按钮点击失败问题
问题描述
代码其余部分运行正常,但点击生成图表的submit按钮与下载CSV的csv-download按钮时出现报错。推测是执行点击命令时,网页对应按钮未处于可视区域,已调整等待时间但无效,寻求解决方法。
原代码
from bs4 import BeautifulSoup from selenium import webdriver from selenium.webdriver.chrome.service import Service from selenium.webdriver.support.ui import Select from selenium.webdriver.chrome import options import unittest import time from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.common.by import By from selenium.webdriver.support import expected_conditions as EC # link to website website = 'https://www.ncei.noaa.gov/access/monitoring/climate-at-a-glance' path = ('../chromedriver') # Folder location where is the chromedriver driver = webdriver.Chrome(path) driver.get(website) driver.maximize_window() # Selection of the sections where the information I am looking for is located state_click = driver.find_element(by=By.XPATH, value='//*[@id="show-statewide"]/a').click() time_series_click = driver.find_element(by=By.XPATH, value='.//*[@id="time-series"]/div[3]/button').click() # selection of the years (for all files the same range of 1950 - 2021) star_year_dropdown =Select(driver.find_element(by=By.ID, value='begyear')) star_year_dropdown.select_by_visible_text('1950') end_year_dropdown = Select(driver.find_element(by=By.ID, value='endyear')) end_year_dropdown.select_by_visible_text('2021') # selection of the parameter to download: Average temperature parameter_dropdown = Select(driver.find_element(by=By.ID, value='parameter')) parameter_dropdown.select_by_visible_text('Average Temperature') # Creating a loop to loop through all the states and all the months: # state selection select_state = driver.find_element(by=By.XPATH, value='.//*[@id="state"]') opcion_state = select_state.find_elements(by=By.TAG_NAME, value='option') # month selection select_month = driver.find_element(by=By.XPATH, value = '//*[@id="month"]') opcion_month = select_month.find_elements(by = By.TAG_NAME, value='option') for option in opcion_month: option.click() for option in opcion_state: option.click() time.sleep(3) plot = driver.find_element(by=By.XPATH, value='.//input[@id="submit"]').click() dowload = WebDriverWait(driver, 10).until(EC.element_to_be_clickable((By.XPATH, '//*[@id="csv-download"]'))).click() time.sleep(3)
解决方法及修改后的代码
核心问题是元素未处于可视区域、循环中元素引用失效以及等待逻辑不够严谨,以下是针对性修复:
关键修改点
- 滚动按钮到可视区域后再点击
- 循环中重新定位元素,避免DOM更新导致的元素引用失效
- 优化等待条件,确保元素完全加载
- 避免循环变量覆盖问题
修改后的完整代码
from bs4 import BeautifulSoup from selenium import webdriver from selenium.webdriver.support.ui import Select import time from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.common.by import By from selenium.webdriver.support import expected_conditions as EC # link to website website = 'https://www.ncei.noaa.gov/access/monitoring/climate-at-a-glance' path = ('../chromedriver') # Folder location where is the chromedriver driver = webdriver.Chrome(path) driver.get(website) driver.maximize_window() # Selection of the sections where the information I am looking for is located driver.find_element(by=By.XPATH, value='//*[@id="show-statewide"]/a').click() driver.find_element(by=By.XPATH, value='.//*[@id="time-series"]/div[3]/button').click() # selection of the years (for all files the same range of 1950 - 2021) star_year_dropdown = Select(driver.find_element(by=By.ID, value='begyear')) star_year_dropdown.select_by_visible_text('1950') end_year_dropdown = Select(driver.find_element(by=By.ID, value='endyear')) end_year_dropdown.select_by_visible_text('2021') # selection of the parameter to download: Average temperature parameter_dropdown = Select(driver.find_element(by=By.ID, value='parameter')) parameter_dropdown.select_by_visible_text('Average Temperature') # Creating a loop to loop through all the states and all the months: for month_idx in range(1, len(Select(driver.find_element(by=By.ID, value='month')).options)): # 重新选择月份,避免DOM更新导致失效 month_dropdown = Select(driver.find_element(by=By.ID, value='month')) month_dropdown.select_by_index(month_idx) for state_idx in range(1, len(Select(driver.find_element(by=By.ID, value='state')).options)): # 重新选择州,避免DOM更新导致失效 state_dropdown = Select(driver.find_element(by=By.ID, value='state')) state_dropdown.select_by_index(state_idx) # 等待submit按钮加载并滚动到可视区域 submit_btn = WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.ID, "submit")) ) driver.execute_script("arguments[0].scrollIntoView(true);", submit_btn) time.sleep(0.5) submit_btn.click() # 等待CSV下载按钮加载、可点击并滚动到可视区域 download_btn = WebDriverWait(driver, 15).until( EC.element_to_be_clickable((By.ID, "csv-download")) ) driver.execute_script("arguments[0].scrollIntoView(true);", download_btn) download_btn.click() time.sleep(3) driver.quit()
额外说明
- 使用
select_by_index替代直接点击<option>元素,更稳定且避免循环中变量覆盖问题 - 移除了未使用的导入(
unittest、Service、options),简化代码 - 添加
driver.quit()确保爬虫结束后关闭浏览器实例
内容的提问来源于stack exchange,提问作者Alejandro
相关产品推荐
相关产品推荐

