Python中如何提取BeautifulSoup解析出的a标签href链接?
解决方案
要提取a标签的href属性值,只需在定位到目标a标签后,通过['href']或get('href')方法获取属性即可。修改后的代码如下:
from bs4 import BeautifulSoup import requests url = 'https://cryptofaucets.cash/2021/02/01/best-free-bitcoin-faucets/' r = requests.get(url).text soup = BeautifulSoup(r,'lxml') l = [] for link in soup.find_all('h2'): # 定位到h2下最后一个a标签 target_a = link.find_all('a')[-1] # 提取href属性值并打印 print(target_a['href']) # 若担心标签无href属性报错,可改用get方法:target_a.get('href')
补充说明
target_a['href']:直接通过键值对形式获取属性,若标签不含href会抛出KeyError。target_a.get('href'):更安全的写法,当属性不存在时返回None而非报错,适合属性存在不确定性的场景。
内容的提问来源于stack exchange,提问作者John Fitzgerald Kennedy
相关产品推荐
相关产品推荐

