You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python Requests无法获取capitoltrades站点多页数据求助

问题分析与解决方案

你遇到的问题大概率是请求头不完整或会话Cookie缺失导致的,以下是具体排查和解决步骤:

1. 补全完整请求头

很多网站会验证User-Agent之外的请求头字段,比如Accept、Accept-Language、Referer等。你可以在Chrome开发者工具(F12)的「网络」面板中,复制浏览器请求该页面时的完整请求头,替换到代码里。

示例修改后的请求头:

headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36',
    'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,image/apng,*/*;q=0.8,application/signed-exchange;v=b3;q=0.7',
    'Accept-Language': 'zh-CN,zh;q=0.9,en;q=0.8',
    'Referer': 'https://www.capitoltrades.com/',
    'Accept-Encoding': 'gzip, deflate, br',
    'Connection': 'keep-alive',
    'Upgrade-Insecure-Requests': '1'
}

2. 先获取会话Cookie再请求

部分网站需要先访问首页建立会话、获取必要Cookie后,才能正常处理带分页参数的请求。修改代码先访问首页:

import bs4 as bs
import requests

payload = {'page': 2, 'per_page': 100}
headers = {
    # 填入上面的完整请求头
}

session = requests.Session()
# 先访问首页获取会话Cookie
session.get('https://www.capitoltrades.com/', headers=headers)
# 再请求目标列表页
r = session.get('https://www.capitoltrades.com/politicians', headers=headers, params=payload)

print(r.url)
soup = bs.BeautifulSoup(r.text, 'html.parser')
politicians = soup.find_all('a', class_='index-card-link')
print(len(politicians))

with open("test.html", "w", encoding='utf-8') as file:
    file.write(soup.prettify())

3. 确认参数传递正确性

先检查print(r.url)输出的URL是否和Chrome中正常工作的URL完全一致,比如是否包含page=2&per_page=100。如果不一致,可能是参数名错误(比如网站实际用page_num而非page),可通过Chrome网络面板查看请求的实际参数名。

4. 排查重定向问题

若网站存在重定向行为,可能会将带参数的请求重定向回首页。可以在get方法中添加allow_redirects=False,查看响应状态码和内容:

r = session.get('https://www.capitoltrades.com/politicians', headers=headers, params=payload, allow_redirects=False)
print(r.status_code)
print(r.text[:500])

内容的提问来源于stack exchange,提问作者JohanW28

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.02 04:31:09