如何使用BeautifulSoup(bs4)去除HTML链接href属性中的无用前缀
decompose() 是BeautifulSoup提供的删除DOM节点的方法,调用会直接移除整个选中的标签,和你的属性修改需求完全不匹配,所以才会出现删除整个a标签的问题。
实现步骤
- 解析HTML内容后定位到目标a标签
- 提取
href属性值,裁剪掉指定前缀后重新赋值给该标签的href属性即可
完整代码示例
from bs4 import BeautifulSoup # 你的原始HTML内容 raw_html = ''' <div class="full-news none"> Demo: <a href="https://www.lolinez.com/?https://www.makemytrip.com" rel="external noopener noreferrer" target="_blank">https://www.makemytrip.com</a> <br/> </div> ''' # 1. 解析HTML soup = BeautifulSoup(raw_html, 'html.parser') # 2. 定位目标a标签,可根据实际场景调整选择器,比如按父类筛选:soup.select_one('.full-news a') target_a = soup.find('a') # 3. 处理href属性,移除指定前缀 prefix = 'https://www.lolinez.com/?' if target_a['href'].startswith(prefix): target_a['href'] = target_a['href'][len(prefix):] # 输出处理后的HTML print(soup.prettify())
效果说明
处理后a标签的href属性会自动更新为https://www.makemytrip.com,其余属性、文本内容和整体HTML结构完全保留,不会影响其他节点。
内容的提问来源于stack exchange,提问作者Sainita
相关产品推荐
相关产品推荐

