You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python使用bs4登录网站后爬取店铺页提示未登录如何解决

问题根源

你代码的核心错误是登录态没有复用:
requests.Session()生成的会话对象c会自动保存登录过程中服务端返回的认证Cookie,所有需要登录状态的请求都必须通过这个会话对象发起。但你请求店铺页面时,直接调用了全局的requests.get(shop_url),这个全新的请求完全不携带任何登录凭证,网站自然判定你处于未登录状态。

修复后可运行代码
import requests
from bs4 import BeautifulSoup

with requests.session() as c: 
    # 加常规请求头模拟真实浏览器,降低被反爬拦截概率
    headers = {
        "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/125.0.0.0 Safari/537.36",
        "Referer": "https://www.tais-shoes.ru/wp-login.php"
    }
    
    link="https://www.tais-shoes.ru/wp-login.php" 
    initial=c.get(link, headers=headers) 

    login_data = {"log": "*****","pwd": "*****", 
              "rememberme": "forever", 
              "redirect_to": "https://www.tais-shoes.ru/my-account/", 
              "redirect_to_automatic": "1"
             }

    page_login = c.post('https://www.tais-shoes.ru/wp-login.php', data=login_data, headers=headers)
    # 验证是否登录成功:如果最终跳转地址是my-account页面说明登录流程正常
    print("登录后跳转地址:", page_login.url)
    print("登录请求状态码:", page_login.status_code)
    
    shop_url = "https://www.tais-shoes.ru/shop/"
    # 核心修复:用保存了登录态的会话对象c发起请求,不要直接调用全局requests.get
    html = c.get(shop_url, headers=headers)
    soup = BeautifulSoup(html.text, 'html.parser')

    print(soup)
排查补充说明
  • 如果改完还是提示未登录,先检查登录请求的返回结果:WordPress站点登录有时会要求携带首次访问登录页时返回的隐藏校验字段(比如安全token、wp-submit字段等),你需要先从initial的响应内容里解析出这些隐藏字段,一起加到login_data里再提交
  • 所有需要登录才能访问的页面请求,都必须放在with requests.session() as c:的代码块内部,通过会话对象c发起,出了with块会话会自动关闭,登录态会直接失效
  • 如果返回403状态码,说明站点有反爬拦截,可以适当增加请求间隔,不要短时间发起大量请求

内容的提问来源于stack exchange,提问作者nasir

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.29 11:54:23