使用BeautifulSoup删除指定类最后一个元素之后的所有HTML内容
实现方法
你不需要用正则匹配字符串处理,直接用BeautifulSoup提供的节点遍历方法就能实现,核心是利用元素的next_siblings属性获取当前元素之后的所有同级节点,逐个删除即可。
完整可运行代码
from bs4 import BeautifulSoup import lxml # 你的HTML示例数据 data = ''' <p>Some text element.</p> <p>Some other text element.</p> <p class="myclass">This is an element with class</p> <p>This is an element without class.</p> <p>Other paragraph.</p> <p class="myclass">The second element with class.</p> <p>Another paragraph.</p> <p>More</p> <p>...</p> ''' soup = BeautifulSoup(data, 'lxml') # 选择所有带指定类的p元素,补全原代码漏掉的右括号 ps_with_class = soup.find_all('p', {'class': 'myclass'}) if ps_with_class: last_p_with_class = ps_with_class[-1] # 遍历最后一个匹配元素之后的所有同级节点,逐个删除 for sibling in list(last_p_with_class.next_siblings): sibling.extract() # 输出处理后的HTML,用body.contents输出可以去掉lxml自动补的html、body标签 print(''.join([str(tag) for tag in soup.body.contents]))
代码说明
next_siblings会返回当前元素之后所有的同级节点,包含HTML标签和文本节点(比如换行、空白字符)- 转成
list遍历是因为直接遍历next_siblings的时候删除节点会导致迭代器出错 extract()方法会把节点从soup树中直接移除,比字符串正则处理更安全,不会误删内容或者破坏HTML结构- 最后输出的时候用
soup.body.contents处理,是因为lxml解析器默认会给HTML补全<html>、<body>根标签,遍历body的子节点输出就能得到你要的原始结构内容。
输出结果
<p>Some text element.</p> <p>Some other text element.</p> <p class="myclass">This is an element with class</p> <p>This is an element without class.</p> <p>Other paragraph.</p> <p class="myclass">The second element with class.</p>
内容的提问来源于stack exchange,提问作者Catalin
相关产品推荐
相关产品推荐

