You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用bs4提取两个strong标签间的所有文本(含无标签内容)

如何用BeautifulSoup提取两个strong标签间的所有文本?

需求:提取<strong>Title 1</strong>之后到<strong>Title 2</strong>之前的所有文本(包括标签包裹和无标签的内容),最终返回格式为 "lorem ipsum 1 lorem ipsum 2 lorem ipsum n"。

待爬取HTML示例:

<p><strong>Title 1</strong>
<br />
lorem ipsum 1</p>
<p>lorem ipsum 2</p>
…
<p>lorem ipsum n</p>

<p><strong>Title 2</strong>
<br />
blah blah </p>

尝试过的代码及问题

第一段代码

# Parse the HTML content using BeautifulSoup
soup = BeautifulSoup(response.content, 'html.parser')

# Find the <strong> tag with the specified text in the section argument
strong_tag = soup.find('strong', string="Title 1")
print("TAG", strong_tag)
if strong_tag:
    # Retrieve all text following the <strong> tag until the next <strong> tag
    section_text = ''
    next_sibling = strong_tag.next_sibling
    print("NEXT SIBLING", next_sibling)
    while next_sibling:
        if next_sibling.string and next_sibling.name != 'strong':
            section_text += next_sibling.string.strip() + ' '
            print("SECTION TEXT", section_text)
            next_sibling = next_sibling.next_sibling
        else:
            break
    
    if not section_text:
        next_tag = strong_tag.find_next()
        print("FIND_NEXT", next_tag)
        while next_tag and next_tag.name != 'strong':
            if next_tag.string:
                print("FIND_NEXT.STRING", next_tag.string)
                section_text += next_tag.string.strip() + ' '
            next_tag = next_tag.find_next()
    
    return section_text.strip()
else:
    print(f"Section '{section}' not found.")
    return None

问题:仅返回 "lorem ipsum 2 lorem ipsum n",缺失 "lorem ipsum 1"。

第二段代码

strong_tag = soup.find('strong', string="Title 1")

if strong_tag:
    # Retrieve all text until the next <strong> tag, regardless of its position
    section_text = ''
    print("TAG", strong_tag)
    while strong_tag:
        if strong_tag.string:
            # Append text
            section_text += strong_tag.string.strip() + ' '
        next_item = strong_tag.next_sibling
        print("NEXTITEM", next_item)
        while next_item and not hasattr(next_item, 'name') and not isinstance(next_item, str):
            # Append text nodes not wrapped in tags
            section_text += next_item.string.strip() + ' '
            next_item = next_item.next_sibling
        if not next_item:
            # Stop if there is no next sibling
            break
        if next_item.name == 'strong':
            # Stop if next tag is a <strong> tag
            break
        strong_tag = next_item

    return section_text.strip()
else:
    print(f"Section '{section}' not found.")
    return None

问题:仅返回 "lorem ipsum 1"。

正确解决方案

核心思路:找到起始strong标签后,遍历所有后续节点(文本节点、标签节点),直到遇到目标结束strong标签,收集所有节点的有效文本并格式化。

from bs4 import BeautifulSoup

def extract_between_strongs(html_content):
    soup = BeautifulSoup(html_content, 'html.parser')
    start_tag = soup.find('strong', string="Title 1")
    if not start_tag:
        print("Section 'Title 1' not found.")
        return None
    
    section_text = []
    current_node = start_tag.next_sibling  # 从起始strong的下一个节点开始遍历
    
    while current_node:
        # 遇到结束的strong标签则停止
        if hasattr(current_node, 'name') and current_node.name == 'strong' and current_node.string == "Title 2":
            break
        
        # 收集文本:纯文本节点直接处理,标签节点提取内部所有文本
        if isinstance(current_node, str):
            text = current_node.strip()
            if text:
                section_text.append(text)
        else:
            tag_text = current_node.get_text(strip=True)
            if tag_text:
                section_text.append(tag_text)
        
        # 移动到下一个节点
        current_node = current_node.next_sibling
    
    # 用空格拼接所有有效文本
    return ' '.join(section_text)

代码说明

  1. 起始定位:精准找到包含Title 1的strong标签,从它的下一个节点开始遍历。
  2. 终止判断:遍历过程中一旦匹配到包含Title 2的strong标签,立即停止。
  3. 文本收集:
    • 纯文本节点:清理首尾空白后,非空文本才加入结果列表。
    • 标签节点:用get_text(strip=True)提取标签内所有层级的文本,同样只保留非空内容。
  4. 结果格式化:将列表中的文本片段用空格连接成最终字符串,符合需求格式。

该方法能覆盖所有层级的文本内容,无论文本是在起始strong的同级节点,还是嵌套在其他标签内,都能完整提取。

内容的提问来源于stack exchange,提问作者Jay Jung

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.24 20:03:12