You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python检测网站是否基于WordPress开发?

检测网站是否基于WordPress开发的改进方案

原来只检查页面里的"wp-content"或"wordpress"没效果,大概率是因为站点做了反爬拦截、修改了静态资源路径,或者这些关键词被隐藏了。可以试试下面这些更可靠的检测方法,结合requests和BeautifulSoup实现:

核心检测思路

  • 检查页面的generator meta标签:大部分WordPress站点会在页面头部标注<meta name="generator" content="WordPress X.X.X">
  • 查看robots.txt文件:里面通常会包含/wp-admin/、wp-includes/这类WordPress特有的路径规则
  • 尝试访问/wp-admin/路径:即使被拦截返回403,页面里也大概率会出现WordPress相关标识
  • 扩大关键词范围:除了wp-content,还可以检查wp-includes这类更难修改的核心路径

代码实现

import requests
from bs4 import BeautifulSoup

def is_wordpress_site(url):
    # 补全URL协议头
    if not url.startswith(('http://', 'https://')):
        url = f'https://{url}'
    
    # 模拟浏览器请求头,避免被反爬拦截
    headers = {
        'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
    }

    try:
        # 1. 检测页面内容
        resp = requests.get(url, headers=headers, timeout=10)
        resp.raise_for_status()
        soup = BeautifulSoup(resp.text, 'html.parser')

        # 检查generator标签
        generator_meta = soup.find('meta', attrs={'name': 'generator'})
        if generator_meta and 'WordPress' in generator_meta.get('content', ''):
            return True

        # 检查核心关键词(转小写避免大小写问题)
        wp_keywords = ['wp-content', 'wp-includes', 'wp-admin', 'wordpress']
        if any(keyword in resp.text.lower() for keyword in wp_keywords):
            return True

        # 2. 检测robots.txt
        robots_url = f"{url.rstrip('/')}/robots.txt"
        robots_resp = requests.get(robots_url, headers=headers, timeout=5)
        if robots_resp.status_code == 200:
            if any(keyword in robots_resp.text.lower() for keyword in wp_keywords):
                return True

        # 3. 检测wp-admin页面
        admin_url = f"{url.rstrip('/')}/wp-admin/"
        admin_resp = requests.get(admin_url, headers=headers, timeout=5)
        # 正常登录页返回200,被拦截返回403,两种情况都可能包含WordPress标识
        if admin_resp.status_code in [200, 403] and ('WordPress' in admin_resp.text or 'wp-login' in admin_resp.text):
            return True

    except requests.exceptions.RequestException:
        # 处理超时、连接失败等异常
        pass

    return False

# 批量检测示例
target_sites = ['wordpress.org', 'example.com', 'your-test-site.com']
for site in target_sites:
    print(f"{site}: {'是WordPress站点' if is_wordpress_site(site) else '不是WordPress站点'}")

注意事项

  • 一定要加User-Agent,不然很多站点会直接拒绝请求
  • 多维度检测能提高准确率,单个特征可能被站点刻意隐藏
  • 异常处理不能少,避免某个站点连接失败导致整个批量检测中断

内容的提问来源于stack exchange,提问作者Puzino

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 10:22:39