如何在BeautifulSoup中删除包含指定域名链接的父级元素
实现代码
from bs4 import BeautifulSoup # 传入待处理的原始HTML raw_html = ''' <pre><code><strong><a href="https://www.fertilizer.com/2021/07/bvfcl.html" target="_blank">Fertilizer Corporation Limited</a> (BVFCL)</strong> has released an employment notification for the recruitment of <strong>11 DGM, Company Secretary, Finance Manager and Accounts Officer Vacancy</strong> </code></pre> ''' # 解析HTML soup = BeautifulSoup(raw_html, 'html.parser') # 筛选所有href属性包含fertilizer.com的a标签 target_a = soup.find_all('a', href=lambda val: val and 'fertilizer.com' in val) for a_node in target_a: # 若只需删除a标签的直接父元素(示例中为strong标签),替换为a_node.parent.decompose()即可 # 要实现示例输出null的效果,直接定位最外层的pre标签删除 pre_parent = a_node.find_parent('pre') if pre_parent: pre_parent.decompose() # 校验剩余内容,为空输出null result = soup.get_text(strip=True) print('null' if not result else soup)
操作说明
- 用
find_all加lambda筛选器可以精准定位所有href带目标域名的a标签,避免误删其他内容 - BeautifulSoup中每个节点都有
.parent属性可以直接获取直接父级,如果要找更上层的指定父标签,用.find_parent(标签名)方法更稳妥,不需要手动数层级 decompose()方法会直接把当前节点和所有子节点从DOM树中彻底移除,不需要额外操作
内容的提问来源于stack exchange,提问作者Sainita
相关产品推荐
相关产品推荐

