如何用Beautiful Soup提取HTML中的href属性链接?
问题
我编写了如下代码,希望从HTML中提取href属性对应的店铺链接(示例链接:https://storelocator.homebargains.co.uk/store/A779/Quedgeley+Retail+Park,+Gloucester),请问该如何实现?
import requests from bs4 import BeautifulSoup url = "https://storelocator.homebargains.co.uk/all-stores" soup = BeautifulSoup(requests.get(url).text, "html.parser") info = soup.find("td") print(info)
解决方案
你的当前代码仅获取了第一个<td>标签,无法提取所有目标链接。以下是两种可行的实现方式:
方式一:精准定位标签提取
先定位包含店铺名称的<td>(页面中这类标签通常带有store-name类),再从中提取<a>标签的href属性:
import requests from bs4 import BeautifulSoup url = "https://storelocator.homebargains.co.uk/all-stores" response = requests.get(url) soup = BeautifulSoup(response.text, "html.parser") # 获取所有带store-name类的td标签 store_cells = soup.find_all("td", class_="store-name") for cell in store_cells: link_tag = cell.find("a") if link_tag and link_tag.get("href"): print(link_tag.get("href"))
方式二:通过链接特征过滤
利用目标链接包含/store/的特征,直接筛选符合条件的<a>标签:
import requests from bs4 import BeautifulSoup url = "https://storelocator.homebargains.co.uk/all-stores" soup = BeautifulSoup(requests.get(url).text, "html.parser") # 遍历所有带href属性的a标签,过滤出包含/store/的链接 for link in soup.find_all("a", href=True): if "/store/" in link["href"]: print(link["href"])
关键说明
- 用
find_all()替代find(),可获取所有匹配标签,而非仅第一个 get("href")比直接["href"]更安全,无href属性时返回None而非报错- 通过class或链接特征过滤,能避免提取无关的冗余链接
内容的提问来源于stack exchange,提问作者S L
相关产品推荐
相关产品推荐

