登录ukwhoswho.com爬取订阅内容时遇Error 403求助排查
问题描述
我需要获取https://www.ukwhoswho.com/的订阅制内容,需通过网站左侧登录框(非右上角「个人资料」入口)提交用户名和密码访问锁定内容,但尝试以下两段Python requests代码均返回Error 403,无法完成登录,请求排查问题:
第一段代码
import requests params = {'user':username, 'pass':password} headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 6.1; WOW64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/56.0.2924.76 Safari/537.36'} # This is chrome, you can set whatever browser you like r = requests.post('https://www.ukwhoswho.com/LOGIN', data=params, headers=headers) r
第二段代码
import requests post_url = 'https://www.ukwhoswho.com/LOGIN' client = requests.session() r = client.get('https://www.ukwhoswho.com/') header_info = { "Host": "www.ukwhoswho.com", "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:107.0) Gecko/20100101 Firefox/107.0", "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8", "Accept-Language": "en-US,en;q=0.5", "Accept-Encoding": "gzip, deflate, br", "Content-Type": "application/x-www-form-urlencoded", "Content-Length": "39", "Origin": "https://www.ukwhoswho.com", "DNT": "1", "Connection": "keep-alive", "Referer": "https://www.ukwhoswho.com/" } payload = {'user':username, 'pass':password} r = client.post(post_url, data=payload, headers = header_info) print(r.text) print(r.status_code)
问题排查与修复方案
导致403的核心原因及修复方式如下:
- 缺失CSRF令牌验证:网站的登录表单大概率包含CSRF令牌(隐藏的input字段),用于防止跨站请求伪造。你的代码未提取并提交该令牌,服务器直接拒绝请求。
- 硬编码Content-Length:第二段代码手动设置了
Content-Length: "39",但实际请求体长度与该值不匹配,服务器判定请求非法,需删除该头,由requests自动计算。 - 冗余请求头干扰:
Host、Connection等头无需手动设置,requests会自动处理,手动设置可能引发格式错误。
修复后的代码示例
import requests from bs4 import BeautifulSoup # 替换为你的账号密码 username = "你的用户名" password = "你的密码" login_url = "https://www.ukwhoswho.com/LOGIN" home_url = "https://www.ukwhoswho.com/" # 初始化会话,自动维护Cookie session = requests.Session() # 访问首页,获取CSRF令牌和会话Cookie home_resp = session.get(home_url) soup = BeautifulSoup(home_resp.text, "html.parser") # 提取左侧登录表单中的CSRF令牌(需根据页面实际源码调整input的name属性) # 示例:查看页面源码,找到类似<input type="hidden" name="csrf_token" value="xxx">的字段 csrf_token = soup.find("input", {"name": "csrf_token"})["value"] # 构造登录请求参数 payload = { "user": username, "pass": password, "csrf_token": csrf_token # 必须加入CSRF令牌 } # 仅保留关键请求头 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:107.0) Gecko/20100101 Firefox/107.0", "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8", "Accept-Language": "en-US,en;q=0.5", "Referer": home_url } # 发送登录请求 login_resp = session.post(login_url, data=payload, headers=headers) # 验证登录结果 if login_resp.status_code == 200: # 登录成功后,使用会话访问订阅内容页面 content_resp = session.get("https://www.ukwhoswho.com/目标订阅页面URL") print(content_resp.text) else: print(f"登录失败,状态码: {login_resp.status_code}") print(login_resp.text)
注意事项
- 需根据网站实际页面源码调整CSRF令牌的
name属性,比如可能是_csrf、token等,可通过浏览器开发者工具查看左侧登录表单的隐藏字段获取。 - 全程使用同一个
Session对象,确保Cookie在请求间自动传递。 - 避免手动设置
Content-Length、Host等系统级请求头,防止请求格式错误。
内容的提问来源于stack exchange,提问作者sepehr
相关产品推荐
相关产品推荐

