求助:使用Python脚本下载SAM排除列表的弹窗激活问题
解决SAM排除列表下载脚本中Accept按钮激活问题
我正在编写Python脚本从SAM网站下载最新排除列表,流程是:访问目标页面→点击初始OK弹窗→点击首个Zip下载链接→处理需要滚动到底部才能激活Accept按钮的弹窗。目前脚本能弹出最后一个窗口,但无法激活Accept按钮完成下载。
修正后的完整脚本
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import os import time def download_latest_file(): try: print("Initializing Chrome WebDriver...") options = webdriver.ChromeOptions() # 设置下载目录,替换为你的实际路径 prefs = {'download.default_directory': '/Users/m_keiffer/Downloads'} options.add_experimental_option('prefs', prefs) driver = webdriver.Chrome(options=options) print("Opening the webpage...") driver.get("https://sam.gov/data-services/Exclusions/Public%20V2?privacy=Public") wait = WebDriverWait(driver, 30) # 处理初始OK弹窗 try: print("Handling initial OK pop-up...") # 匹配包含"OK"文本的按钮 ok_button = wait.until(EC.element_to_be_clickable((By.XPATH, "//button[contains(text(), 'OK')]"))) ok_button.click() print("OK button clicked.") except Exception as e: print(f"Error handling initial OK pop-up: {str(e)}") # 找到并点击首个Zip下载链接 try: print("Finding download links...") # 定位页面中第一个Zip格式的下载链接 exclusions_link = wait.until(EC.element_to_be_clickable((By.XPATH, "//a[contains(@href, '.zip')][1]"))) print(f"Clicking on the download link: {exclusions_link.text}") exclusions_link.click() print("Download link clicked.") except Exception as e: print(f"Error finding or clicking download link: {str(e)}") driver.quit() return # 处理Accept弹窗 try: print("Handling Accept pop-up...") # 等待弹窗加载完成,定位弹窗内的内容滚动区域 modal_content = wait.until(EC.presence_of_element_located((By.XPATH, "//sa-security-modal//div[@class='modal-body']"))) # 滚动到内容区域底部,触发Accept按钮激活逻辑 driver.execute_script("arguments[0].scrollTop = arguments[0].scrollHeight", modal_content) # 等待按钮状态更新,替换sleep为显式等待更可靠 accept_button = wait.until(EC.element_to_be_clickable((By.XPATH, "//button[contains(text(), 'Accept')]"))) accept_button.click() print("Accept button clicked.") except Exception as e: print(f"Error handling Accept pop-up: {str(e)}") driver.quit() return # 等待下载完成(可根据文件大小调整时间,或通过文件状态判断) time.sleep(15) download_path = "/Users/m_keiffer/Downloads" files_after = os.listdir(download_path) print(f"Files in Downloads: {files_after}") new_files = [f for f in files_after if f.endswith('.zip') or f.endswith('.csv')] if new_files: # 按文件修改时间排序获取最新文件 latest_file = max(new_files, key=lambda x: os.path.getmtime(os.path.join(download_path, x))) print(f"Latest downloaded file: {latest_file}") else: print("No new files found in the Downloads folder.") driver.quit() except Exception as e: print(f"Failed to initialize the WebDriver or other error: {str(e)}") # 执行函数 download_latest_file()
关键修改点说明
- 补全并修正XPath定位:原脚本中XPath被截断,替换为精准匹配元素的路径,确保能正确找到OK按钮、下载链接和Accept按钮。
- 滚动正确的元素:针对弹窗内的
modal-body区域滚动,而非整个弹窗容器,确保触发页面的滚动检测逻辑来激活Accept按钮。 - 替换sleep为显式等待:使用
WebDriverWait等待Accept按钮可点击,比固定sleep更可靠,避免因加载延迟导致的失败。 - 完善最新文件判断逻辑:原脚本中
max函数的key参数不完整,补充为按文件修改时间排序,精准获取最新下载的文件。
内容的提问来源于stack exchange,提问作者Michael Keiffer
相关产品推荐
相关产品推荐

