如何用Python检测网站是否基于WordPress开发?
检测网站是否基于WordPress开发的改进方案
原来只检查页面里的"wp-content"或"wordpress"没效果,大概率是因为站点做了反爬拦截、修改了静态资源路径,或者这些关键词被隐藏了。可以试试下面这些更可靠的检测方法,结合requests和BeautifulSoup实现:
核心检测思路
- 检查页面的
generatormeta标签:大部分WordPress站点会在页面头部标注<meta name="generator" content="WordPress X.X.X"> - 查看
robots.txt文件:里面通常会包含/wp-admin/、wp-includes/这类WordPress特有的路径规则 - 尝试访问/wp-admin/路径:即使被拦截返回403,页面里也大概率会出现WordPress相关标识
- 扩大关键词范围:除了wp-content,还可以检查wp-includes这类更难修改的核心路径
代码实现
import requests from bs4 import BeautifulSoup def is_wordpress_site(url): # 补全URL协议头 if not url.startswith(('http://', 'https://')): url = f'https://{url}' # 模拟浏览器请求头,避免被反爬拦截 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36' } try: # 1. 检测页面内容 resp = requests.get(url, headers=headers, timeout=10) resp.raise_for_status() soup = BeautifulSoup(resp.text, 'html.parser') # 检查generator标签 generator_meta = soup.find('meta', attrs={'name': 'generator'}) if generator_meta and 'WordPress' in generator_meta.get('content', ''): return True # 检查核心关键词(转小写避免大小写问题) wp_keywords = ['wp-content', 'wp-includes', 'wp-admin', 'wordpress'] if any(keyword in resp.text.lower() for keyword in wp_keywords): return True # 2. 检测robots.txt robots_url = f"{url.rstrip('/')}/robots.txt" robots_resp = requests.get(robots_url, headers=headers, timeout=5) if robots_resp.status_code == 200: if any(keyword in robots_resp.text.lower() for keyword in wp_keywords): return True # 3. 检测wp-admin页面 admin_url = f"{url.rstrip('/')}/wp-admin/" admin_resp = requests.get(admin_url, headers=headers, timeout=5) # 正常登录页返回200,被拦截返回403,两种情况都可能包含WordPress标识 if admin_resp.status_code in [200, 403] and ('WordPress' in admin_resp.text or 'wp-login' in admin_resp.text): return True except requests.exceptions.RequestException: # 处理超时、连接失败等异常 pass return False # 批量检测示例 target_sites = ['wordpress.org', 'example.com', 'your-test-site.com'] for site in target_sites: print(f"{site}: {'是WordPress站点' if is_wordpress_site(site) else '不是WordPress站点'}")
注意事项
- 一定要加User-Agent,不然很多站点会直接拒绝请求
- 多维度检测能提高准确率,单个特征可能被站点刻意隐藏
- 异常处理不能少,避免某个站点连接失败导致整个批量检测中断
内容的提问来源于stack exchange,提问作者Puzino
相关产品推荐
相关产品推荐

