You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用BeautifulSoup框架获取section标签的所有子元素内容?

获取section标签内部内容的几种方法

针对你的需求,以下是几种用BeautifulSoup提取<section class="post-content">内部内容的实用方法:

方法1:获取直接子节点(含文本节点)

find_all返回匹配元素的列表,先取出目标section元素,再用.contents获取它的所有直接子节点(包括换行、空格这类文本节点):

from bs4 import BeautifulSoup

html = '''
<body>
  <section class="post-content">
    <h1>title</h1>
    <div>balabala</div>
  </section>
<body>
'''
soup = BeautifulSoup(html, 'html.parser')
# 取出第一个匹配的section元素
section = soup.find_all("section", {"class": "post-content"})[0]
# 获取内部直接子节点
inner_contents = section.contents
print(inner_contents)

输出会包含所有直接子节点,包括空白文本节点。

方法2:获取内部HTML字符串(保留标签结构)

如果需要保留内部的HTML标签结构,用.decode_contents()方法,返回section内部的HTML代码:

inner_html = section.decode_contents()
print(inner_html)

输出结果:

<h1>title</h1>
    <div>balabala</div>

方法3:提取所有纯文本内容(自动去空白)

如果只需要提取内部的纯文本,去掉多余换行和空格,用.stripped_strings迭代器拼接成字符串:

text = ' '.join(section.stripped_strings)
print(text)

输出:title balabala

方法4:仅获取子标签元素(排除文本节点)

如果只想获取section下的标签元素,排除空白文本节点,有两种实现方式:

# 方式A:用find_all直接获取所有子标签
inner_tags = section.find_all(True)
print(inner_tags)  # 输出[<h1>title</h1>, <div>balabala</div>]

# 方式B:遍历children过滤非标签节点
inner_tags = [child for child in section.children if child.name is not None]
print(inner_tags)

内容的提问来源于stack exchange,提问作者蔡忠振

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 12:40:32