如何使用BeautifulSoup提取网页首个class为clear的div之前的所有元素
解决方法
核心逻辑为:先定位到第一个<div class="clear"></div>分隔节点,仅收集该节点之前的现役站点a标签即可。
修改后可直接运行的代码
import requests from bs4 import BeautifulSoup # 替换为你自己的请求头配置 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36' } hhref = ['https://www.ukfirestations.co.uk/stations/bedfordshire','https://www.ukfirestations.co.uk/stations/buckinghamshire'] dats = [] for url in hhref: r = requests.get(url, headers=headers) soup = BeautifulSoup(r.content, 'html.parser') stations_grid = soup.find('div', id='stations-grid') # 定位第一个分隔现役/停用站点的clear标签 first_clear_tag = stations_grid.find('div', class_='clear') # 获取分隔标签前所有a标签,reverse=True保证和网页顺序一致 active_stations = first_clear_tag.find_all_previous('a', href=True, reverse=True) for station in active_stations: dats.append(station['href'])
可选的遍历实现方案
如果需要更灵活的自定义过滤规则,也可以选择遍历父容器子节点,遇到分隔标签直接终止:
stations_grid = soup.find('div', id='stations-grid') for child in stations_grid.children: # 匹配到第一个clear标签直接停止遍历 if child.name == 'div' and 'clear' in child.get('class', []): break # 仅收集带href属性的有效a标签 if child.name == 'a' and child.get('href'): dats.append(child['href'])
内容的提问来源于stack exchange,提问作者Stackcans
相关产品推荐
相关产品推荐

