使用Python Requests.get请求Investing.com数据时遇403 Forbidden错误求助
解决requests请求Investing.com数据返回403的问题
问题原因分析
不一定是JWT导致,但核心是网站的反爬机制验证不通过:
- 浏览器请求时会自动携带完整的会话Cookie、动态生成的请求参数(比如你URL里的随机串),而脚本直接硬写旧参数/缺失Cookie,被服务器判定为非合法请求。
- curl/Postman请求失败也是因为同样的原因——没有还原浏览器请求的全部上下文。
实用解决建议
1. 携带完整会话Cookie
打开浏览器开发者工具(F12),找到历史数据请求的Cookie头,把所有Cookie内容复制到脚本的headers里:
import requests session = requests.Session() # 先访问主页面获取初始Cookie(保证会话一致性) session.get("https://www.investing.com/indices/nifty-50") headers = { # 新增Cookie,替换为浏览器复制的完整内容 'Cookie': '这里替换成浏览器里复制的完整Cookie字符串', 'authority': 'tvc4.investing.com', 'accept': '*/*', 'accept-language': 'en-US,en;q=0.9', 'content-type': 'text/plain', 'origin': 'https://tvc-invdn-com.investing.com', 'referer': 'https://tvc-invdn-com.investing.com/', 'sec-ch-ua': '"Chromium";v="116", "Not)A;Brand";v="24", "Google Chrome";v="116"', 'sec-ch-ua-mobile': '?0', 'sec-ch-ua-platform': '"Windows"', 'sec-fetch-dest': 'empty', 'sec-fetch-mode': 'cors', 'sec-fetch-site': 'same-site', 'user-agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/116.0.0.0 Safari/537.36', } params = { 'symbol': '17940', 'resolution': '1', 'from': '1694613203', 'to': '1694699663', } # 替换URL中的动态串为浏览器最新请求的内容 response = session.get( 'https://tvc4.investing.com/[最新动态串]/[最新时间戳]/56/56/23/history', params=params, headers=headers, ) print(response.status_code) print(response.text)
2. 替换URL中的动态参数
你当前URL里的86fd2cdaadcb4c77ffd862dcab14c196和1694699591是动态生成的(可能是请求令牌或时间戳),每次请求都会变化:
- 临时解决:刷新浏览器,从开发者工具复制最新的历史数据请求URL替换到脚本中。
- 长期解决:分析主页面的JS代码,找到这些参数的生成逻辑,实现自动生成。
3. 用Session保持会话一致性
使用requests.Session()代替直接调用requests.get(),它会自动保存和复用Cookie,模拟浏览器的会话过程,避免手动管理Cookie的麻烦。
4. 最后考虑模拟浏览器(Selenium/Playwright)
如果以上方法都无法绕过反爬,再用Selenium或Playwright模拟真实浏览器行为——它们会自动处理Cookie、动态参数、JS渲染等问题,但缺点是速度较慢、资源占用高。示例(Selenium):
from selenium import webdriver from selenium.webdriver.common.by import By import time driver = webdriver.Chrome() driver.get("https://www.investing.com/indices/nifty-50-historical-data") # 等待页面加载完成 time.sleep(3) # 获取历史数据表格内容(需根据页面实际元素调整) table = driver.find_element(By.ID, "curr_table") rows = table.find_elements(By.TAG_NAME, "tr") for row in rows[1:]: # 跳过表头行 print(row.text) driver.quit()
内容的提问来源于stack exchange,提问作者Junior2691
相关产品推荐
相关产品推荐

