使用Python爬取NSE 52周新高数据表遇阻,求技术支持
解决NSE 52周新高/新低数据爬取问题
方法一:直接下载CSV文件(最便捷)
NSE提供了直接获取数据的API接口,无需爬取页面,直接请求即可拿到CSV格式数据:
import requests # 52周新高数据接口 url_high = "https://www.nseindia.com/api/52-week-high-low?index=equities" # 52周新低数据接口 url_low = "https://www.nseindia.com/api/52-week-high-low?index=equities&type=low" # 添加请求头,避免被拦截 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' } # 下载并保存新高数据 response_high = requests.get(url_high, headers=headers) with open('nse_52week_high.csv', 'wb') as f: f.write(response_high.content) # 下载并保存新低数据 response_low = requests.get(url_low, headers=headers) with open('nse_52week_low.csv', 'wb') as f: f.write(response_low.content)
方法二:用Selenium+BeautifulSoup爬取页面表格
如果需要从页面爬取表格,需等待页面动态加载完成后再解析,修正后的代码如下:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from bs4 import BeautifulSoup import pandas as pd # 初始化Chrome浏览器(需确保ChromeDriver与Chrome版本匹配) driver = webdriver.Chrome() driver.get("https://www.nseindia.com/market-data/52-week-high-low") # 等待表格加载完成(最多等待10秒) wait = WebDriverWait(driver, 10) wait.until(EC.presence_of_element_located((By.CLASS_NAME, 'equity-table'))) # 解析页面源码中的表格 soup = BeautifulSoup(driver.page_source, 'html.parser') table = soup.find('table', class_='equity-table') # 转换为DataFrame并输出 df = pd.read_html(str(table))[0] print(df.head()) # 保存为CSV文件 df.to_csv('nse_52week_table.csv', index=False) # 关闭浏览器 driver.quit()
注意事项
- 确保ChromeDriver版本与本地Chrome浏览器版本一致,否则会启动失败
- 可添加
--headless=new参数启用无头模式,避免弹出浏览器窗口:from selenium.webdriver.chrome.options import Options options = Options() options.add_argument('--headless=new') driver = webdriver.Chrome(options=options)
内容的提问来源于stack exchange,提问作者Jaya Vamshi
相关产品推荐
相关产品推荐

