Python BeautifulSoup:如何获取停止标签之前的目标标签?
使用BeautifulSoup获取class="nice"标签并排除stop之后的内容
方法一:利用兄弟元素定位
先找到stop标签,再获取它之前所有同级的nice标签,注意find_previous_siblings返回的是倒序结果,需要反转过来:
from bs4 import BeautifulSoup html = ''' <div class="nice"></div> <div class="nice"></div> <div class="stop">here should be the end of found items</div> <div class="nice"></div> <div class="nice"></div> ''' soup = BeautifulSoup(html, 'html.parser') # 定位stop标签 stop_tag = soup.find('div', class_='stop') # 获取stop之前的所有同级nice标签 nice_tags = stop_tag.find_previous_siblings('div', class_='nice') # 反转列表恢复正序 nice_tags = list(reversed(nice_tags)) # 输出结果 for tag in nice_tags: print(tag)
方法二:遍历元素直到遇到stop终止
遍历目标区域的元素,收集nice标签,一旦碰到stop标签就停止遍历,这种方式更灵活,适合复杂层级结构:
from bs4 import BeautifulSoup html = ''' <div class="nice"></div> <div class="nice"></div> <div class="stop">here should be the end of found items</div> <div class="nice"></div> <div class="nice"></div> ''' soup = BeautifulSoup(html, 'html.parser') nice_tags = [] # 遍历所有直接子元素(如果目标div在某个父容器里,替换成对应的父元素即可) for tag in soup.body.children: # 跳过文本节点(比如换行符) if not tag.name: continue if tag.name == 'div': # 遇到stop标签就终止遍历 if 'stop' in tag.get('class', []): break # 收集nice标签 if 'nice' in tag.get('class', []): nice_tags.append(tag) # 输出结果 for tag in nice_tags: print(tag)
两种方法都能得到stop标签之前的两个nice标签,根据你的实际HTML结构选择即可。
内容的提问来源于stack exchange,提问作者Fred
相关产品推荐
相关产品推荐

