如何仅从soup.find而非find_all的结果中获取所有href属性值
从指定父节点提取所有匹配href属性的解决方法
你当前的写法核心问题在于findAll返回的是符合条件的标签列表,不存在.item属性,且你已经拿到了match_items父节点对象,直接在该对象上调用find_all就只会在当前节点内部搜索,不会触发全局搜索,完全符合你的需求。
修正后的代码如下:
import requests from bs4 import BeautifulSoup url_news = "https://www.hltv.org/matches" response = requests.get(url_news) soup = BeautifulSoup(response.content, "html.parser") match_info = [] # 拿到父节点对象 match_items = soup.find("div", class_="upcomingMatchesSection") # 增加非空判断避免页面结构变更时报错 if match_items: # 仅在match_items节点内部搜索符合条件的a标签,不会全局搜索 match_a_list = match_items.find_all("a", class_="match a-reset", href=True) # 遍历所有匹配标签提取href存入列表 for a_tag in match_a_list: match_info.append(a_tag["href"])
补充说明:
- BeautifulSoup的所有元素查找方法都支持在任意标签节点上调用,调用范围仅限该节点的后代节点,你不需要调用soup全局的
find_all就能完成局部查找。 - 如果你要更精简的写法也可以用列表推导式:
match_info = [a["href"] for a in match_items.find_all("a", class_="match a-reset", href=True)],前提是确认match_items不为空。
内容的提问来源于stack exchange,提问作者makim
相关产品推荐
相关产品推荐

