You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python3使用BS4筛选并打印包含.php?id=的指定href链接

代码修改方案

错误原因

  • 直接使用通配符*.php?id=*做包含判断不符合Python语法,in关键字无法识别通配符规则
  • 未遍历单个href链接逐个校验,直接对整个href列表做匹配,不可能命中目标

基于你现有lxml实现的修改代码

import requests
from lxml import html

def gethref(ip):
    url = f"http://{ip}"
    print(f"[x] ~ SCAN: {url} ~ [x]")
    # 增加请求异常捕获,避免站点无法访问时报错退出
    try:
        req = requests.get(url, timeout=10)
        # 自动适配页面编码,避免乱码
        req.encoding = req.apparent_encoding
        tree = html.fromstring(req.text)
        # 提取所有href属性
        all_href = tree.xpath('//@href')
        # 遍历筛选符合条件的链接
        for href in all_href:
            # 先判断href不为空,再匹配子串
            if href and '.php?id=' in href:
                print(href)
    except Exception as e:
        print(f"访问站点出错:{str(e)}")

基于BeautifulSoup的实现代码(对应你注释里的方案)

import requests
from bs4 import BeautifulSoup

def gethref(ip):
    url = f"http://{ip}"
    print(f"[x] ~ SCAN: {url} ~ [x]")
    try:
        req = requests.get(url, timeout=10)
        req.encoding = req.apparent_encoding
        soup = BeautifulSoup(req.text, 'html.parser')
        # 查找所有a标签
        for a_tag in soup.find_all('a'):
            href = a_tag.get('href')
            if href and '.php?id=' in href:
                print(href)
    except Exception as e:
        print(f"访问站点出错:{str(e)}")

可选优化

如果需要更精准的匹配规则,比如仅匹配id参数为数字的链接,可以使用正则表达式做过滤:

  1. 先导入re库:import re
  2. 把判断条件替换为:if href and re.search(r'\.php\?id=\d+', href)

内容的提问来源于stack exchange,提问作者00xZ

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.06 18:54:04