You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用BeautifulSoup提取网页首个class为clear的div之前的所有元素

解决方法

核心逻辑为:先定位到第一个<div class="clear"></div>分隔节点,仅收集该节点之前的现役站点a标签即可。

修改后可直接运行的代码

import requests
from bs4 import BeautifulSoup

# 替换为你自己的请求头配置
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
}
hhref = ['https://www.ukfirestations.co.uk/stations/bedfordshire','https://www.ukfirestations.co.uk/stations/buckinghamshire']

dats = []
for url in hhref:
    r = requests.get(url, headers=headers)
    soup = BeautifulSoup(r.content, 'html.parser')
    stations_grid = soup.find('div', id='stations-grid')
    # 定位第一个分隔现役/停用站点的clear标签
    first_clear_tag = stations_grid.find('div', class_='clear')
    # 获取分隔标签前所有a标签,reverse=True保证和网页顺序一致
    active_stations = first_clear_tag.find_all_previous('a', href=True, reverse=True)
    for station in active_stations:
        dats.append(station['href'])

可选的遍历实现方案

如果需要更灵活的自定义过滤规则,也可以选择遍历父容器子节点,遇到分隔标签直接终止:

stations_grid = soup.find('div', id='stations-grid')
for child in stations_grid.children:
    # 匹配到第一个clear标签直接停止遍历
    if child.name == 'div' and 'clear' in child.get('class', []):
        break
    # 仅收集带href属性的有效a标签
    if child.name == 'a' and child.get('href'):
        dats.append(child['href'])

内容的提问来源于stack exchange,提问作者Stackcans

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.07 03:18:03