如何用Python获取网站上动态更新的zip文件链接?
解决ERCOT动态加载文件下载的方案
方案一:用Selenium模拟浏览器抓取
BeautifulSoup只能解析静态HTML,动态渲染的内容得等JS执行完毕才能获取,用Selenium可以模拟浏览器加载全过程,完美解决这个问题:
- 先安装依赖:
pip install selenium,再下载对应浏览器的驱动(比如ChromeDriver,需和浏览器版本匹配) - 示例代码:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC # 初始化Chrome浏览器 driver = webdriver.Chrome() driver.get("https://www.ercot.com/mp/data-products/compliance-and-disclosure/?id=NP3-965-ER") # 等待文件列表加载完成,最多等待10秒 try: # 定位第一个包含.zip的下载链接,可根据页面实际元素调整xpath download_link = WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.XPATH, "//a[contains(text(), 'Download') and contains(@href, '.zip')]")) ) zip_url = download_link.get_attribute("href") print(f"最新ZIP文件链接: {zip_url}") # 直接下载文件 import requests response = requests.get(zip_url) with open("latest_ercot_data.zip", "wb") as f: f.write(response.content) finally: driver.quit()
方案二:抓包调用后台API(更高效)
不用启动浏览器,直接找到网站加载文件列表的后台接口:
- 打开浏览器F12,切换到「网络」标签页
- 刷新目标页面,过滤「XHR」或「Fetch」请求,找到返回文件列表的接口
- 查看该接口的请求参数和返回的JSON结构,里面会包含带动态doclookup id的下载链接
- 用requests直接调用接口解析数据,示例代码:
import requests # 替换为你抓包找到的API地址 api_url = "https://www.ercot.com/api/xxx/xxx" headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } response = requests.get(api_url, headers=headers) data = response.json() # 按发布时间排序取最新的zip文件 latest_doc = sorted(data["documents"], key=lambda x: x["publishDate"], reverse=True)[0] zip_url = latest_doc["downloadUrl"] print(f"最新ZIP文件链接: {zip_url}")
额外提示
- 可以查下ERCOT是否有官方数据API文档,官方接口比抓包更稳定
- 用Selenium时建议开启无头模式(添加
options.add_argument("--headless=new")),减少资源占用 - 注意控制请求频率,避免触发网站反爬机制
内容的提问来源于stack exchange,提问作者Reza Tabrizi
相关产品推荐
相关产品推荐

