You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python爬虫检测WordPress站点触发403 Forbidden错误如何解决

问题原因分析
  • 核心错误:你配置的请求头没有作用到实际抓取内容的请求上。你代码中先用requests.get()带自定义头发起了一次请求,但后续真正读取页面内容用的是urllib.request.urlopen(url),这个调用没有携带任何自定义请求头,相当于裸请求访问,自然会被目标站点的安全策略拦截返回403。
  • 额外问题:两种HTTP请求库混用完全没有必要,平白增加出错概率。
修复方案

优先推荐统一使用requests库,修正后的代码如下:

import requests

def check_web_wp(url):
    is_wordpress = False
    headers = {
        "User-Agent": "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36",
        "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8",
        "Accept-Language": "fr-FR,fr;q=0.8,en-US;q=0.5,en;q=0.3",
        "Accept-Encoding": "gzip, deflate, br",
        "DNT": "1",
        "Connection": "keep-alive",
        "Upgrade-Insecure-Requests": "1",
        "Sec-Fetch-Dest": "document",
        "Sec-Fetch-Mode": "navigate",
        "Sec-Fetch-Site": "none",
        "Sec-Fetch-User": "?1"
    }
    try:
        response = requests.get(url, headers=headers, timeout=10)
        response.raise_for_status()
        # 直接检查响应文本,拆分单词会额外消耗性能
        if "wordpress" in response.text.lower() or "wp-content" in response.text or "wp-includes" in response.text:
            is_wordpress = True
    except requests.exceptions.RequestException as e:
        print(f"请求出错:{e}")
    return is_wordpress


def main():
    url = "https://icalendrier.fr/"
    is_wp = check_web_wp(url)
    print(f"站点是否为WordPress:{is_wp}")

if __name__ == "__main__":
    main()
补充优化建议
  • 不要仅靠页面关键词判断WordPress,可直接请求/wp-admin、/wp-content等WordPress默认路径,检测返回状态可以大幅提升准确率和检测效率
  • 新增的Sec-Fetch-*系列请求头,更贴近真实浏览器的请求特征,降低被拦截概率
  • 增加了异常捕获逻辑,避免请求超时、访问错误直接导致脚本崩溃
  • 更新了较新版本的Chrome User-Agent,你原本使用的Firefox 66版本UA已经非常老旧,容易被安全策略识别为爬虫

内容的提问来源于stack exchange,提问作者pgmendormi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.04 13:15:03