You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何修改Python爬虫代码仅抓取指定层级(level-2)的href链接,避免获取下层(level-3)元素

解决BeautifulSoup抓取指定层级href的问题

我明白你的问题了——你在用BeautifulSoup爬取欧洲杯2020的分组赛页面时,原本只想抓取level-2层级下的目标href链接,但代码不小心把下层level-3的链接也收集进来了。别担心,这是因为默认的findAll会递归查找所有后代元素,我们只需要调整选择器,精准定位level-2的直接子元素就行。

问题根源

你原来的代码里,soup.find('ul', class_='level-2').findAll('li')会匹配ul.level-2下所有的li元素,包括嵌套在里面的level-3层级的li,这就是为什么会多出不需要的链接。

修改后的解决方案

这里有两种简单的方法可以实现精准抓取:

方法1:使用recursive=False限制查找范围

通过给findAll添加recursive=False参数,让它只查找当前节点的直接子元素,不递归深入后代:

import bs4 as bs
import requests

url = 'https://int.soccerway.com/international/europe/european-championships/2020/group-stage/r38188/'
headers = {"User-agent":"Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/77.0.3865.120 Safari/537.36"}
resp = requests.get(url, headers=headers)
soup = bs.BeautifulSoup(resp.text, 'lxml')

# 仅获取level-2 ul的直接子li元素
ls = soup.find('ul', class_='level-2').findAll('li', recursive=False)
for i in ls:
    print(i.find('a')['href'])

方法2:使用CSS子选择器(更简洁直观)

CSS的>子选择器可以直接定位父元素的直接子节点,用soup.select实现会更简洁:

import bs4 as bs
import requests

url = 'https://int.soccerway.com/international/europe/european-championships/2020/group-stage/r38188/'
headers = {"User-agent":"Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/77.0.3865.120 Safari/537.36"}
resp = requests.get(url, headers=headers)
soup = bs.BeautifulSoup(resp.text, 'lxml')

# 用CSS子选择器直接匹配level-2 ul的子li
ls = soup.select('ul.level-2 > li')
for i in ls:
    print(i.find('a')['href'])

预期输出

修改后运行代码,就能得到你想要的两个精准链接:

  • /international/europe/european-championships/2020/group-stage/r38188/
  • /international/europe/european-championships/2020/s13030/final-stages/

内容的提问来源于stack exchange,提问作者Resultados Oficiais

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 03:19:08