如何用Python登录Finviz网站并自动下载CSV文件?
问题描述
我尝试从finviz.com导出CSV文件,手动通过浏览器登录并导出可以成功,但通过Python代码实现却无法正常运行。以下是我的代码:
import requests login_url = 'https://www.finviz.com/login_submit.ashx' file_url = 'https://elite.finviz.com/export.ashx?v=152&c=1,,65,66,67,25,64,63,49' payload = { 'email': 'email@gmail.com', 'password': 'myPassword', 'remember': 'true'} headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/58.0.3029.110 Safari/537.3' } with requests.Session() as s: s.post(login_url, headers=headers, data=payload) response = s.get(file_url) if response.ok: with open('file.csv', 'wb') as f: f.write(response.content) else: print(f'Failed to download the file: status code {response.status_code}')
注意:如果直接使用浏览器认证后的headers(包含Cookie),代码可以正常运行,但每次Cookie过期都需要重新生成。以下是可用的代码:
import requests login_url = 'https://www.finviz.com/login_submit.ashx' file_url = 'https://elite.finviz.com/export.ashx?v=152&c=1,,65,66,67,25,64,63,49' payload = { 'email': 'email@gmail.com', 'password': 'myPassword', 'remember': 'true'} headers = { 'authority': 'www.google-analytics.com', 'accept': '*/*', 'accept-language': 'en-US,en;q=0.9', 'cookie': '<long_cookie_string>', 'sec-ch-ua': '"Not.A/Brand";v="8", "Chromium";v="114", "Google Chrome";v="114"', 'sec-ch-ua-mobile': '?0', 'sec-ch-ua-platform': '"Windows"', 'sec-fetch-dest': 'empty', 'sec-fetch-mode': 'no-cors', 'sec-fetch-site': 'cross-site', 'sec-fetch-user': '?1', 'upgrade-insecure-requests': '1', 'user-agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36', 'Referer': 'https://elite.finviz.com/screener.ashx?v=152&c=1,,65,66,67,25,64,63,49', 'Origin': 'https://elite.finviz.com', 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36', 'Accept': '*/*', 'Accept-Language': 'en-US,en;q=0.9', 'Connection': 'keep-alive', 'Content-Type': 'text/plain;charset=UTF-8', 'Sec-Fetch-Dest': 'empty', 'Sec-Fetch-Mode': 'cors', 'Sec-Fetch-Site': 'cross-site', 'content-type': 'application/x-www-form-urlencoded', 'origin': 'https://elite.finviz.com', 'referer': 'https://elite.finviz.com/', 'content-length': '0', } with requests.Session() as s: response = requests.get(file_url, headers=headers) if response.ok: with open('file.csv', 'wb') as f: f.write(response.content) else: print(f'Failed to download the file: status code {response.status_code}')
请问我哪里出错了?如何用Python创建会话并管理Cookie?
问题分析与解决方案
你的代码存在的问题
- 缺少初始会话Cookie:Finviz的登录接口需要先获取服务器生成的初始会话Cookie(比如会话标识、CSRF令牌),直接调用登录提交接口会因为缺少这些Cookie导致登录请求被拒绝,会话未正确建立。
- 请求头不完整:仅携带User-Agent不足以模拟浏览器请求,Finviz可能校验Referer、Origin等关键头信息,识别出自动化请求后拒绝服务。
- 未验证登录结果:你直接发送登录请求但未检查是否成功,若登录失败(比如账号错误、请求被拦截),后续下载请求自然无法通过认证。
正确的会话与Cookie管理方法
核心步骤
- 先访问登录页面获取初始Cookie:通过GET请求访问Finviz登录页,让服务器生成初始会话Cookie,
requests.Session会自动保存这些Cookie。 - 完善请求头:添加浏览器登录时的关键头信息,让请求更接近正常浏览器行为。
- 校验登录状态:发送登录请求后,通过响应状态码或跳转结果确认登录是否成功,避免后续请求基于未认证会话执行。
修正后的代码
import requests # 定义关键URL login_page_url = 'https://www.finviz.com/login.ashx' login_submit_url = 'https://www.finviz.com/login_submit.ashx' file_url = 'https://elite.finviz.com/export.ashx?v=152&c=1,,65,66,67,25,64,63,49' # 登录参数 payload = { 'email': 'email@gmail.com', 'password': 'myPassword', 'remember': 'true' } # 模拟浏览器的完整请求头 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36', 'Referer': 'https://www.finviz.com/login.ashx', 'Origin': 'https://www.finviz.com' } with requests.Session() as s: # 1. 获取初始会话Cookie s.get(login_page_url, headers=headers) # 2. 发送登录请求并校验结果 login_response = s.post(login_submit_url, headers=headers, data=payload) if login_response.status_code != 200 or 'login' in login_response.url: print('登录失败,请检查账号密码或请求配置') exit() # 3. 下载CSV文件 file_response = s.get(file_url, headers=headers) if file_response.ok: with open('file.csv', 'wb') as f: f.write(file_response.content) print('文件下载成功') else: print(f'文件下载失败,状态码:{file_response.status_code}')
额外优化:Cookie持久化
如果不想每次运行都重新登录,可以将会话Cookie保存到文件,下次直接加载:
import pickle # 保存Cookie到文件 with open('finviz_cookies.pkl', 'wb') as f: pickle.dump(s.cookies, f) # 下次运行时加载Cookie with requests.Session() as s: with open('finviz_cookies.pkl', 'rb') as f: s.cookies.update(pickle.load(f)) # 直接尝试下载文件 file_response = s.get(file_url, headers=headers) # 后续逻辑同上
内容的提问来源于stack exchange,提问作者cttrader
相关产品推荐
相关产品推荐

