如何用bs4提取两个strong标签间的所有文本(含无标签内容)
如何用BeautifulSoup提取两个strong标签间的所有文本?
需求:提取<strong>Title 1</strong>之后到<strong>Title 2</strong>之前的所有文本(包括标签包裹和无标签的内容),最终返回格式为 "lorem ipsum 1 lorem ipsum 2 lorem ipsum n"。
待爬取HTML示例:
<p><strong>Title 1</strong> <br /> lorem ipsum 1</p> <p>lorem ipsum 2</p> … <p>lorem ipsum n</p> <p><strong>Title 2</strong> <br /> blah blah </p>
尝试过的代码及问题
第一段代码
# Parse the HTML content using BeautifulSoup soup = BeautifulSoup(response.content, 'html.parser') # Find the <strong> tag with the specified text in the section argument strong_tag = soup.find('strong', string="Title 1") print("TAG", strong_tag) if strong_tag: # Retrieve all text following the <strong> tag until the next <strong> tag section_text = '' next_sibling = strong_tag.next_sibling print("NEXT SIBLING", next_sibling) while next_sibling: if next_sibling.string and next_sibling.name != 'strong': section_text += next_sibling.string.strip() + ' ' print("SECTION TEXT", section_text) next_sibling = next_sibling.next_sibling else: break if not section_text: next_tag = strong_tag.find_next() print("FIND_NEXT", next_tag) while next_tag and next_tag.name != 'strong': if next_tag.string: print("FIND_NEXT.STRING", next_tag.string) section_text += next_tag.string.strip() + ' ' next_tag = next_tag.find_next() return section_text.strip() else: print(f"Section '{section}' not found.") return None
问题:仅返回 "lorem ipsum 2 lorem ipsum n",缺失 "lorem ipsum 1"。
第二段代码
strong_tag = soup.find('strong', string="Title 1") if strong_tag: # Retrieve all text until the next <strong> tag, regardless of its position section_text = '' print("TAG", strong_tag) while strong_tag: if strong_tag.string: # Append text section_text += strong_tag.string.strip() + ' ' next_item = strong_tag.next_sibling print("NEXTITEM", next_item) while next_item and not hasattr(next_item, 'name') and not isinstance(next_item, str): # Append text nodes not wrapped in tags section_text += next_item.string.strip() + ' ' next_item = next_item.next_sibling if not next_item: # Stop if there is no next sibling break if next_item.name == 'strong': # Stop if next tag is a <strong> tag break strong_tag = next_item return section_text.strip() else: print(f"Section '{section}' not found.") return None
问题:仅返回 "lorem ipsum 1"。
正确解决方案
核心思路:找到起始strong标签后,遍历所有后续节点(文本节点、标签节点),直到遇到目标结束strong标签,收集所有节点的有效文本并格式化。
from bs4 import BeautifulSoup def extract_between_strongs(html_content): soup = BeautifulSoup(html_content, 'html.parser') start_tag = soup.find('strong', string="Title 1") if not start_tag: print("Section 'Title 1' not found.") return None section_text = [] current_node = start_tag.next_sibling # 从起始strong的下一个节点开始遍历 while current_node: # 遇到结束的strong标签则停止 if hasattr(current_node, 'name') and current_node.name == 'strong' and current_node.string == "Title 2": break # 收集文本:纯文本节点直接处理,标签节点提取内部所有文本 if isinstance(current_node, str): text = current_node.strip() if text: section_text.append(text) else: tag_text = current_node.get_text(strip=True) if tag_text: section_text.append(tag_text) # 移动到下一个节点 current_node = current_node.next_sibling # 用空格拼接所有有效文本 return ' '.join(section_text)
代码说明
- 起始定位:精准找到包含
Title 1的strong标签,从它的下一个节点开始遍历。 - 终止判断:遍历过程中一旦匹配到包含
Title 2的strong标签,立即停止。 - 文本收集:
- 纯文本节点:清理首尾空白后,非空文本才加入结果列表。
- 标签节点:用
get_text(strip=True)提取标签内所有层级的文本,同样只保留非空内容。
- 结果格式化:将列表中的文本片段用空格连接成最终字符串,符合需求格式。
该方法能覆盖所有层级的文本内容,无论文本是在起始strong的同级节点,还是嵌套在其他标签内,都能完整提取。
内容的提问来源于stack exchange,提问作者Jay Jung
相关产品推荐
相关产品推荐

