使用Python Selenium下载NSE CSV文件时遇ERR_HTTP2_PROTOCOL_ERROR
解决NSE India内幕交易CSV下载的ERR_HTTP2_PROTOCOL_ERROR问题
问题背景
我尝试使用以下Python Selenium脚本,从NSE India下载4个月时间范围的内幕交易CSV文件并重命名为Insider.CSV:
from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.chrome.options import Options import os import time def download_csv_with_selenium(): # 设置Chrome选项指定下载目录和请求头 chrome_options = Options() download_directory = os.path.dirname(os.path.abspath("Insider.csv")) chrome_options.add_experimental_option("prefs", {"download.default_directory": download_directory}) chrome_options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/88.0.4324.182 Safari/537.36") chrome_options.add_argument("accept-language=en-US,en;q=0.9") chrome_options.add_argument("accept-encoding=gzip, deflate, br") chrome_options.add_argument("referer=https://www.nseindia.com/") # 初始化Chrome驱动 driver = webdriver.Chrome(options=chrome_options) try: # 打开内幕交易页面 driver.get("https://www.nseindia.com/companies-listing/corporate-filings-insider-trading") # 滚动页面 driver.execute_script("window.scrollBy(0, 300);") # 点击3M时间范围按钮 button_3m = WebDriverWait(driver, 120).until( EC.element_to_be_clickable((By.XPATH, "//button[text()='3M']")) ) button_3m.click() time.sleep(10) # 点击下载CSV按钮 button_download = WebDriverWait(driver, 120).until( EC.element_to_be_clickable((By.XPATH, "//button[text()='Download(.csv)']")) ) button_download.click() time.sleep(10) # 重命名下载文件 downloaded_file_path = os.path.join(download_directory, "DownloadedFile.csv") os.rename(downloaded_file_path, "Insider.csv") print("File downloaded and renamed successfully: Insider.csv") finally: driver.quit() if __name__ == "__main__": download_csv_with_selenium()
运行时出现ERR_HTTP2_PROTOCOL_ERROR,提示API链接可能临时失效或永久迁移,无法完成下载。
解决方案
1. 禁用Chrome的HTTP/2协议
NSE服务器对HTTP/2兼容性不佳,在ChromeOptions中添加禁用参数:
chrome_options.add_argument("--disable-http2")
2. 更新用户代理字符串
旧版本UA可能被服务器拦截,替换为现代Chrome的UA:
chrome_options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36")
3. 替换固定等待为文件下载监听
不要用time.sleep()等待下载完成,改用目录监听确保文件下载完成:
import glob def wait_for_download(download_dir, timeout=30): start_time = time.time() while time.time() - start_time < timeout: # 筛选已完成的CSV文件(排除Chrome临时下载文件) completed_files = [ f for f in glob.glob(os.path.join(download_dir, "*.csv")) if not f.endswith(".crdownload") ] if completed_files: return completed_files[0] time.sleep(1) raise TimeoutError("下载超时") # 在点击下载按钮后调用 downloaded_file = wait_for_download(download_directory) os.rename(downloaded_file, os.path.join(download_directory, "Insider.csv"))
4. 添加更多合规请求头
补充必要的请求头,模拟真实浏览器行为:
chrome_options.add_argument("accept=text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8") chrome_options.add_argument("upgrade-insecure-requests=1")
5. 直接调用API(替代Selenium方案)
如果Selenium仍有问题,改用requests库直接请求API,效率更高:
import requests headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36", "Referer": "https://www.nseindia.com/companies-listing/corporate-filings-insider-trading", "Accept-Language": "en-US,en;q=0.9", "Accept-Encoding": "gzip, deflate, br", "Accept": "text/csv,*/*;q=0.9" } url = "https://www.nseindia.com/api/corporates-pit?index=equities&from_date=24-09-2023&to_date=24-12-2023&csv=true" response = requests.get(url, headers=headers) response.raise_for_status() # 抛出HTTP错误 with open("Insider.csv", "wb") as f: f.write(response.content) print("文件下载完成:Insider.csv")
注意事项
- NSE的API参数和验证规则可能随时变更,需定期检查
- 频繁请求可能触发反爬限制,建议添加合理的请求间隔
- 直接调用API若遇403错误,可先请求NSE主页获取会话Cookie后再调用API
内容的提问来源于stack exchange,提问作者nikunj baheti
相关产品推荐
相关产品推荐

