爬取FINRA数据时如何获取XSRF Token?是否存在动态生成无法获取的情况?
从FINRA网站获取数据时的XSRF Token获取问题
我尝试从FINRA网站获取数据并转换为JSON格式,再存入DataFrame。参考Stack Overflow的方法,试图在请求中提交XSRF Token,但未找到页面中存储该Token的隐藏输入字段,只能硬编码从其他方案获取的Token,但这类Token会过期。我通过requests.Session和BeautifulSoup检查了请求的Cookie、Header及页面元素,均未找到该Token,因此疑问:
- 该Token是否由JavaScript动态生成导致无法获取?
- 我是否遗漏了某些步骤?
以下是我的代码:
import requests import json import pandas as pd from bs4 import BeautifulSoup import requests session = requests.Session() print(session.cookies.get_dict()) response = session.get('https://www.finra.org/finra-data/fixed-income/corp-and-agency') print("Cookies: ", session.cookies) # Nothing with the token print("Param: ", session.params) # Nothing with the token print("Header: ", session.params) # Nothing with the token get_token_response = requests.get('https://www.finra.org/finra-data/fixed-income/corp-and-agency') soup = BeautifulSoup(get_token_response.text) print(get_token_response.headers) # Nothing with the token print(get_token_response.cookies) # Nothing with the token print(soup.findAll('input')) # 修正原代码缺失的括号 # Hard Coded headers = { 'authority': 'services-dynarep.ddwa.finra.org', 'accept': 'application/json, text/plain, */*', 'content-type': 'application/json', 'cookie':'XSRF-TOKEN=578706e6-5dfa-4beb-b887-60b42da068be;', 'origin': 'https://www.finra.org', 'referer': 'https://www.finra.org/', 'x-xsrf-token': '578706e6-5dfa-4beb-b887-60b42da068be', } data = ('{"fields":["issueSymbolIdentifier","issuerName","isCallable","productSubTypeCode",' '"couponRate","maturityDate","industryGroup","moodysRating",' '"standardAndPoorsRating","lastSalePrice","lastSaleYield"],' '"dateRangeFilters":[],"domainFilters":[],"compareFilters":[],' '"multiFieldMatchFilters":[{"fuzzy":false,"searchValue":"gme","synonym":true,"fields":' '[{"name":"issuerName","boost":1}]}],"orFilters":[],"aggregationFilter":null,' '"sortFields":["+issuerName"],"limit":50,"offset":0,"delimiter":null,"quoteValues":false}') response = requests.post('https://services-dynarep.ddwa.finra.org/public/reporting/v2/data/group/FixedIncomeMarket/name/CorporateAndAgencySecurities', headers=headers, data=data) print(response.status_code) data = json.dumps(response.json()['returnBody']['data'], indent=4) data = list(data.replace('\n', '').replace('\'', '')) data = ''.join(data[1:-1]) df = pd.read_json(data)
解决方案
1. XSRF Token的生成方式
你的猜测正确,这个XSRF Token确实是JavaScript动态生成的,不会直接出现在页面HTML或初始请求的Cookie/Header里。FINRA的前端会在页面加载后通过JS生成Token,并存入Cookie,后续请求需要带上这个Token才能通过验证。
2. 正确获取Token的步骤
你之前的代码没有利用Session的连贯性,且分开发起请求导致无法捕获JS生成的Cookie。下面是两种可行方案:
方案一:使用支持JS渲染的工具(如Playwright)
这类工具可以模拟浏览器执行JS,直接获取生成的Token:
from playwright.sync_api import sync_playwright import requests import pandas as pd with sync_playwright() as p: # 启动无头浏览器 browser = p.chromium.launch(headless=True) page = browser.new_page() # 访问目标页面触发JS生成Token page.goto('https://www.finra.org/finra-data/fixed-income/corp-and-agency') # 从Cookie中提取XSRF-TOKEN cookies = page.context.cookies() xsrf_token = next((cookie['value'] for cookie in cookies if cookie['name'] == 'XSRF-TOKEN'), None) browser.close() # 用获取到的Token发起请求 session = requests.Session() session.cookies.set('XSRF-TOKEN', xsrf_token) headers = { 'authority': 'services-dynarep.ddwa.finra.org', 'accept': 'application/json, text/plain, */*', 'content-type': 'application/json', 'origin': 'https://www.finra.org', 'referer': 'https://www.finra.org/finra-data/fixed-income/corp-and-agency', 'x-xsrf-token': xsrf_token, } # 用字典格式定义请求数据,避免手动拼接JSON data = { "fields": ["issueSymbolIdentifier","issuerName","isCallable","productSubTypeCode", "couponRate","maturityDate","industryGroup","moodysRating", "standardAndPoorsRating","lastSalePrice","lastSaleYield"], "dateRangeFilters": [], "domainFilters": [], "compareFilters": [], "multiFieldMatchFilters": [{"fuzzy": False,"searchValue": "gme","synonym": True,"fields": [{"name": "issuerName","boost": 1}]}], "orFilters": [], "aggregationFilter": None, "sortFields": ["+issuerName"], "limit": 50, "offset": 0, "delimiter": None, "quoteValues": False } # 发送POST请求,直接用json参数提交数据 response = session.post( 'https://services-dynarep.ddwa.finra.org/public/reporting/v2/data/group/FixedIncomeMarket/name/CorporateAndAgencySecurities', headers=headers, json=data ) # 直接将返回的JSON数据转换为DataFrame df = pd.DataFrame(response.json()['returnBody']['data']) print(df.head())
方案二:查找并调用Token生成接口
通过浏览器开发者工具的Network面板,筛选XHR/fetch请求,找到FINRA网站生成XSRF Token的专门接口,直接用requests调用该接口获取Token。这种方式无需模拟浏览器,但需要你自行定位接口。
3. 代码优化建议
- 避免重复导入
requests模块 - 提交请求数据时使用
json=data参数,无需手动拼接JSON字符串,减少出错概率 - 解析返回数据时,直接用
pd.DataFrame转换response.json()['returnBody']['data'],无需多余的字符串处理
内容的提问来源于stack exchange,提问作者David Frick
相关产品推荐
相关产品推荐

