You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

登录ukwhoswho.com爬取订阅内容时遇Error 403求助排查

问题描述

我需要获取https://www.ukwhoswho.com/的订阅制内容,需通过网站左侧登录框(非右上角「个人资料」入口)提交用户名和密码访问锁定内容,但尝试以下两段Python requests代码均返回Error 403,无法完成登录,请求排查问题:

第一段代码

import requests
params = {'user':username, 'pass':password}
headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 6.1; WOW64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/56.0.2924.76 Safari/537.36'} # This is chrome, you can set whatever browser you like
r = requests.post('https://www.ukwhoswho.com/LOGIN', data=params, headers=headers)
r

第二段代码

import requests

post_url = 'https://www.ukwhoswho.com/LOGIN'

client = requests.session()
r = client.get('https://www.ukwhoswho.com/')

header_info = {
"Host": "www.ukwhoswho.com",
"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:107.0) Gecko/20100101 Firefox/107.0",
"Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8",
"Accept-Language": "en-US,en;q=0.5",
"Accept-Encoding": "gzip, deflate, br",
"Content-Type": "application/x-www-form-urlencoded",
"Content-Length": "39",
"Origin": "https://www.ukwhoswho.com",
"DNT": "1",
"Connection": "keep-alive",
"Referer": "https://www.ukwhoswho.com/"
}

payload = {'user':username, 'pass':password} 

r = client.post(post_url, data=payload, headers = header_info)

print(r.text)
print(r.status_code)
问题排查与修复方案

导致403的核心原因及修复方式如下:

  • 缺失CSRF令牌验证:网站的登录表单大概率包含CSRF令牌(隐藏的input字段),用于防止跨站请求伪造。你的代码未提取并提交该令牌,服务器直接拒绝请求。
  • 硬编码Content-Length:第二段代码手动设置了Content-Length: "39",但实际请求体长度与该值不匹配,服务器判定请求非法,需删除该头,由requests自动计算。
  • 冗余请求头干扰:Host、Connection等头无需手动设置,requests会自动处理,手动设置可能引发格式错误。

修复后的代码示例

import requests
from bs4 import BeautifulSoup

# 替换为你的账号密码
username = "你的用户名"
password = "你的密码"

login_url = "https://www.ukwhoswho.com/LOGIN"
home_url = "https://www.ukwhoswho.com/"

# 初始化会话,自动维护Cookie
session = requests.Session()

# 访问首页,获取CSRF令牌和会话Cookie
home_resp = session.get(home_url)
soup = BeautifulSoup(home_resp.text, "html.parser")

# 提取左侧登录表单中的CSRF令牌(需根据页面实际源码调整input的name属性)
# 示例:查看页面源码,找到类似<input type="hidden" name="csrf_token" value="xxx">的字段
csrf_token = soup.find("input", {"name": "csrf_token"})["value"]

# 构造登录请求参数
payload = {
    "user": username,
    "pass": password,
    "csrf_token": csrf_token  # 必须加入CSRF令牌
}

# 仅保留关键请求头
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:107.0) Gecko/20100101 Firefox/107.0",
    "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8",
    "Accept-Language": "en-US,en;q=0.5",
    "Referer": home_url
}

# 发送登录请求
login_resp = session.post(login_url, data=payload, headers=headers)

# 验证登录结果
if login_resp.status_code == 200:
    # 登录成功后,使用会话访问订阅内容页面
    content_resp = session.get("https://www.ukwhoswho.com/目标订阅页面URL")
    print(content_resp.text)
else:
    print(f"登录失败,状态码: {login_resp.status_code}")
    print(login_resp.text)

注意事项

  • 需根据网站实际页面源码调整CSRF令牌的name属性,比如可能是_csrf、token等,可通过浏览器开发者工具查看左侧登录表单的隐藏字段获取。
  • 全程使用同一个Session对象,确保Cookie在请求间自动传递。
  • 避免手动设置Content-Length、Host等系统级请求头,防止请求格式错误。

内容的提问来源于stack exchange,提问作者sepehr

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 05:50:29