Python无Cookie头POST请求超时问题及股指数据自动化获取方案
问题概述
想要自动化从Nifty指数官网获取股指历史数据,目前通过requests.post调用接口时必须手动设置完整请求头和Cookie:
- 移除任意请求头(如Cookie)会触发
ReadTimeout错误 - Cookie随会话变更,无法长期复用
- 手动维护大量请求头不利于自动化流程
当前可用但存在缺陷的代码:
import requests url = "https://www.niftyindices.com/Backpage.aspx/getHistoricaldatatabletoString" json_payload = {'name': 'NIFTY AUTO', 'startDate': '01-Feb-2023', 'endDate': '01-Feb-2024'} headers = { 'Accept': 'application/json, text/javascript, */*; q=0.01', 'Accept-Encoding': 'gzip, deflate, br, zstd', 'Accept-Language': 'en-GB,en;q=0.9', 'Connection': 'keep-alive', 'Content-Type': 'application/json; charset=UTF-8', 'Cookie': 'AbCd1234', 'Host': 'www.niftyindices.com', 'Origin': 'https://www.niftyindices.com', 'Referer': 'https://www.niftyindices.com/reports/historical-data', 'Sec-Fetch-Dest': 'empty', 'Sec-Fetch-Mode': 'cors', 'Sec-Fetch-Site': 'same-origin', 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/122.0.0.0 Safari/537.36', 'X-Requested-With': 'XMLHttpRequest', 'sec-ch-ua': '"Chromium";v="122", "Not(A:Brand";v="24", "Google Chrome";v="122"', 'sec-ch-ua-mobile': '?0', 'sec-ch-ua-platform': '"Windows"' } response = requests.post(url, json=json_payload, headers=headers, timeout=30) print(response.status_code) print(response.json())
解决方案
1. 自动管理Cookie:使用requests.Session
Session对象会自动保存和复用会话Cookie,无需手动复制粘贴。先访问历史数据页面获取初始Cookie,再发送POST请求:
import requests # 初始化会话,自动处理Cookie session = requests.Session() # 设置基础请求头,模拟浏览器 base_headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/122.0.0.0 Safari/537.36', 'Accept-Language': 'en-GB,en;q=0.9' } session.headers.update(base_headers) # 先访问历史数据页面,获取会话所需Cookie history_page_url = "https://www.niftyindices.com/reports/historical-data" session.get(history_page_url, timeout=30) # 发送POST请求获取数据 api_url = "https://www.niftyindices.com/Backpage.aspx/getHistoricaldatatabletoString" payload = {'name': 'NIFTY AUTO', 'startDate': '01-Feb-2023', 'endDate': '01-Feb-2024'} # 仅保留服务器校验的必要请求头 post_headers = { 'Accept': 'application/json, text/javascript, */*; q=0.01', 'Content-Type': 'application/json; charset=UTF-8', 'Referer': history_page_url, 'X-Requested-With': 'XMLHttpRequest' } response = session.post(api_url, json=payload, headers=post_headers, timeout=30) print(response.status_code) print(response.json())
2. 精简请求头:保留核心字段即可
大部分浏览器自动生成的请求头(如Host、Connection、Accept-Encoding)会由requests自动处理,只需保留服务器明确校验的字段:
User-Agent:避免被识别为非浏览器请求Accept:指定接受JSON格式响应Content-Type:告知服务器请求体为JSONReferer:模拟从历史数据页面发起请求的场景X-Requested-With:标记为AJAX请求,部分ASP.NET后端会校验此参数
如果精简后仍出现超时,可以尝试添加sec-ch-ua系列头,但多数情况下上述核心字段足够。
额外优化建议
- 若遇到反爬拦截,可使用
fake_useragent库随机生成User-Agent,避免固定标识 - 在请求之间添加1-2秒延迟(
time.sleep(1)),降低请求频率 - 检查页面是否包含ASP.NET特有的
__VIEWSTATE、__VIEWSTATEGENERATOR等参数,若有需从历史页面HTML中提取并加入请求
总结
- 用
requests.Session自动管理Cookie,彻底告别手动复制 - 精简请求头至核心校验字段,减少维护成本
内容的提问来源于stack exchange,提问作者tintin98
相关产品推荐
相关产品推荐

