使用Python+Selenium下载NSE印度站ETF日报失败,页面持续加载无法下载
问题原因
NSE India平台内置了反自动化检测机制,能够识别Selenium启动的浏览器携带的webdriver特征,因此拦截了页面数据加载请求,导致你用Selenium访问时页面一直卡在加载状态,手动访问无特征所以正常。此外你的代码还存在两处语法错误,也会导致执行失败。
修复步骤
- 给Chrome启动参数添加反检测配置,隐藏自动化特征
- 修正显式等待的参数格式,expected_conditions的定位方法要求传入(By定位方式、路径)组成的元组
- 补全缺失的time模块导入
- 把已废弃的find_element_by_xpath写法替换为官方推荐的通用写法
修正后完整代码
import time from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC options = webdriver.ChromeOptions() # 反爬核心配置 options.add_argument("--disable-blink-features=AutomationControlled") options.add_experimental_option("excludeSwitches", ["enable-automation"]) options.add_experimental_option("useAutomationExtension", False) prefs = {"download.default_directory": "/Volumes/Project/WebScraper/downloadData"} options.binary_location = r'/Applications/Google Chrome 2.app/Contents/MacOS/Google Chrome' chrome_driver_binary = r'/usr/local/Caskroom/chromedriver/94.0.4606.61/chromedriver' options.add_experimental_option("prefs", prefs) driver = webdriver.Chrome(chrome_driver_binary, options=options) # 覆盖webdriver特征 driver.execute_cdp_cmd("Page.addScriptToEvaluateOnNewDocument", { "source": "Object.defineProperty(navigator, 'webdriver', {get: () => undefined})" }) try: driver.get('https://www.nseindia.com/market-data/exchange-traded-funds-etf') # 修正显式等待的参数格式,加一层括号组成元组 element = WebDriverWait(driver, 50).until(EC.visibility_of_element_located((By.XPATH, "//table[@id='etfTable']"))) # 替换已废弃的定位写法 downloadcsv = driver.find_element(By.XPATH, "//div[@id='esw-etf']/div[2]/div/div[3]/div/ul/li/a") print(downloadcsv) downloadcsv.click() time.sleep(5) driver.close() except Exception as e: print(f"执行失败,错误信息:{e}")
内容的提问来源于stack exchange,提问作者Rove sprite
相关产品推荐
相关产品推荐

