Python3使用BS4筛选并打印包含.php?id=的指定href链接
代码修改方案
错误原因
- 直接使用通配符
*.php?id=*做包含判断不符合Python语法,in关键字无法识别通配符规则 - 未遍历单个href链接逐个校验,直接对整个href列表做匹配,不可能命中目标
基于你现有lxml实现的修改代码
import requests from lxml import html def gethref(ip): url = f"http://{ip}" print(f"[x] ~ SCAN: {url} ~ [x]") # 增加请求异常捕获,避免站点无法访问时报错退出 try: req = requests.get(url, timeout=10) # 自动适配页面编码,避免乱码 req.encoding = req.apparent_encoding tree = html.fromstring(req.text) # 提取所有href属性 all_href = tree.xpath('//@href') # 遍历筛选符合条件的链接 for href in all_href: # 先判断href不为空,再匹配子串 if href and '.php?id=' in href: print(href) except Exception as e: print(f"访问站点出错:{str(e)}")
基于BeautifulSoup的实现代码(对应你注释里的方案)
import requests from bs4 import BeautifulSoup def gethref(ip): url = f"http://{ip}" print(f"[x] ~ SCAN: {url} ~ [x]") try: req = requests.get(url, timeout=10) req.encoding = req.apparent_encoding soup = BeautifulSoup(req.text, 'html.parser') # 查找所有a标签 for a_tag in soup.find_all('a'): href = a_tag.get('href') if href and '.php?id=' in href: print(href) except Exception as e: print(f"访问站点出错:{str(e)}")
可选优化
如果需要更精准的匹配规则,比如仅匹配id参数为数字的链接,可以使用正则表达式做过滤:
- 先导入re库:
import re - 把判断条件替换为:
if href and re.search(r'\.php\?id=\d+', href)
内容的提问来源于stack exchange,提问作者00xZ
相关产品推荐
相关产品推荐

