Python BeautifulSoup4如何从find_all结果中提取a标签的href链接
问题根源
你遍历得到的x是外层<div>标签的解析对象,并非需要提取属性的<a>标签对象,且select是调用方法需要用括号而非方括号,两者是你之前尝试失败的核心原因。
可行解决方案
方案一:直接定位目标a标签(推荐,效率更高)
不需要先遍历外层div,直接通过特征筛选Github对应的a标签:
import requests from bs4 import BeautifulSoup as bs # 建议加请求头避免被反爬拦截 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/125.0.0.0 Safari/537.36" } coin_link = requests.get('https://www.coingecko.com/en/coins/bitcoin', headers=headers) soup2 = bs(coin_link.text, "html.parser") # 筛选包含Github文本的a标签 github_tag = soup2.find("a", string=lambda t: t and "Github" in t.strip()) if github_tag: github_url = github_tag["href"] print(github_url)
方案二:基于现有代码逻辑修改
在你已拿到的div对象中,进一步查找内部的a标签提取链接:
import requests from bs4 import BeautifulSoup as bs import time headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/125.0.0.0 Safari/537.36" } coin_link = requests.get('https://www.coingecko.com/en/coins/bitcoin', headers=headers) soup2 = bs(coin_link.text, "html.parser") to_find_github_link = soup2.find_all('div',{'class':'tw-flex flex-wrap tw-font-normal'}) for x in to_find_github_link: time.sleep(1) # 从当前div中查找所有带href属性的a标签 a_list = x.find_all("a", href=True) for a in a_list: # 筛选包含github域名的链接 if "github.com" in a["href"]: print(a["href"]) break
内容的提问来源于stack exchange,提问作者Danny
相关产品推荐
相关产品推荐

