无法从Niftyindices.com爬取NIFTY 100历史指数数据求助
问题描述
此前我可从印度NSE旗下的Niftyindices网站爬取所需的历史指数数据,历史数据位于首页的Reports板块下的Historical Data页面。需选择Index Type为Equity、Index为NIFTY 100,数据范围通常选2001年1月1日至当日。截至昨日我仍能通过现有代码获取数据,但今日起相同代码突然无法返回数据,收到的JSON为空。根据过往经验bm_sv是必须传递的Cookie,恳请协助解决问题。
以下是我使用的爬虫代码:
import requests import json from datetime import datetime print('start') data=[] headers = {'Content-Type': 'application/json; charset=utf-8' ,'Accept':'application/json, text/javascript, */*; q=0.01' ,'Accept-Encoding':'gzip, deflate' ,'Accept-Language':'en-US,en;q=0.9' ,'Content-Length':'100' ,'User-Agent':'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/88.0.4324.190 Safari/537.36' ,'Cookie':'bm_sv=9B706239B47F50CA0B651E20BA5CBF74~YAAQFjkgF2zWSiyOAQAAdV0dQRfVSICkZc20SfI+PnDk8taK1Ppu1ZSmjclFkHqVgsGOE0vK3WnPMHuhY5kOStjVm4OnN1wm9SBRO3nIAvXWAVCR8iN23B8R7kHpcme82M8ytCrJ/LozntCxQlQSFqzuFwLw4+ZPBjdkICfQH4piCmjvZB3AH8NvCmf+nbzT34Q4JO4zYeYadkjlKjVRVIh0lzX2BK8crljTE9W+F1DUdtZYBRBUCM83OIfmZhnH6PnDu79C~1' } row=['indexname','date','open','high','low','close'] data.append(row) payload={'name': 'NIFTY 100','startDate':'01-Jan-2001','endDate': '10-Mar-2024'} JSonURL='https://www.niftyindices.com/Backpage.aspx/getHistoricaldatatabletoString' r=requests.post(JSonURL, data=json.dumps(payload),headers=headers) print(r.text) text=r.json() print(text) datab=json.loads(text['d']) sorted_data=sorted(datab,key=lambda x: datetime.strptime(x['HistoricalDate'], '%d %b %Y'), reverse=False) print('startdata available from: ',datetime.strftime(datetime.strptime(sorted_data[0]['HistoricalDate'], '%d %b %Y'),'%d-%b-%Y')) print('data available till',datetime.strftime(datetime.strptime(sorted_data[len(datab)-1]['HistoricalDate'], '%d %b %Y'),'%d-%b-%Y\n')) for rec in sorted_data: row=[] row.append(rec['Index Name']) row.append(datetime.strptime(rec['HistoricalDate'], '%d %b %Y')) row.append(rec['OPEN'].replace('-','0')) row.append(rec['HIGH'].replace('-','0')) row.append(rec['LOW'].replace('-','0')) row.append(rec['CLOSE']) print(row) data.append(row) print(data)
解决方案
问题核心是硬编码的bm_sv Cookie已过期,这类会话Cookie通常有有效期,过期后服务器会拒绝请求。以下是修复步骤:
动态获取最新Cookie
不要直接写死Cookie,先发送GET请求访问网站首页,获取服务器返回的最新Cookie(包括bm_sv),再用这个Cookie发送POST请求。调整请求头参数
- 移除硬编码的
Content-Length,requests会自动计算并添加该字段,手动设置可能和实际请求体长度不匹配导致错误。 - 更新User-Agent为当前主流浏览器版本,降低被识别为爬虫的概率。
- 移除硬编码的
修改后的代码示例
import requests import json from datetime import datetime print('start') data = [] row = ['indexname', 'date', 'open', 'high', 'low', 'close'] data.append(row) # 初始化会话,自动管理Cookie session = requests.Session() session.headers.update({ 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/122.0.0.0 Safari/537.36', 'Accept': 'application/json, text/javascript, */*; q=0.01', 'Accept-Language': 'en-US,en;q=0.9', 'Content-Type': 'application/json; charset=utf-8' }) # 先访问首页获取最新Cookie session.get('https://www.niftyindices.com') payload = {'name': 'NIFTY 100', 'startDate': '01-Jan-2001', 'endDate': '10-Mar-2024'} JSonURL = 'https://www.niftyindices.com/Backpage.aspx/getHistoricaldatatabletoString' r = session.post(JSonURL, data=json.dumps(payload)) print(r.text) text = r.json() if not text.get('d'): print("未获取到数据,可能需要检查请求参数或Cookie有效性") else: datab = json.loads(text['d']) sorted_data = sorted(datab, key=lambda x: datetime.strptime(x['HistoricalDate'], '%d %b %Y'), reverse=False) print('数据起始日期: ', datetime.strftime(datetime.strptime(sorted_data[0]['HistoricalDate'], '%d %b %Y'), '%d-%b-%Y')) print('数据结束日期: ', datetime.strftime(datetime.strptime(sorted_data[-1]['HistoricalDate'], '%d %b %Y'), '%d-%b-%Y\n')) for rec in sorted_data: row = [ rec['Index Name'], datetime.strptime(rec['HistoricalDate'], '%d %b %Y'), rec['OPEN'].replace('-', '0'), rec['HIGH'].replace('-', '0'), rec['LOW'].replace('-', '0'), rec['CLOSE'] ] print(row) data.append(row) print(data)
- 额外注意事项
如果仍然无法获取数据,可通过浏览器开发者工具查看正常请求的头信息,补充Referer等缺失的请求头;若会话方式仍失效,可尝试添加短暂延迟,或先访问历史数据页面再提交请求,模拟真实用户的操作流程。
内容的提问来源于stack exchange,提问作者Kiran Jain
相关产品推荐
相关产品推荐

