如何用Python BeautifulSoup(BS4)提取无标准标签包裹的HREF
解决方案
你的目标HTML存在格式不规范问题(包含href的标签写法错误),但BeautifulSoup依然能解析它的属性。可以通过以下步骤提取第一个href对应的URL:
- 先获取目标
item标签集合(你已完成这一步):
items = soup.find_all('item', class_="sale-item")
- 取第一个
item标签:
if items: first_item = items[0]
- 在这个
item标签下,查找任意带有href属性的子元素(无需在意标签名,适配不规范的HTML场景):
target_element = first_item.find(attrs={"href": True})
- 提取href属性值:
if target_element: url = target_element.get('href') print(url) # 输出: http://www.assus.com/12165456ALPHA.html
简化写法
若能确保目标元素存在,可合并为一行代码:
url = soup.find('item', class_="sale-item").find(attrs={"href": True}).get('href')
核心逻辑是find(attrs={"href": True})会匹配所有包含href属性的元素,完全适配这种标签写法不规范的特殊场景。
内容的提问来源于stack exchange,提问作者walksonair
相关产品推荐
相关产品推荐

