You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何爬取纯HTML长文章中按strong标签分章节的段落并存储

实现方案

你可以直接用Python的BeautifulSoup库完成提取,操作步骤如下:

  1. 先安装依赖库
pip install beautifulsoup4
  1. 运行以下代码即可输出你需要的变量内容:
from bs4 import BeautifulSoup

# 替换为你自己的原始HTML长文本
raw_html = '''
<div id="book">
<h3>This is book</h3> 

<p> <strong> Chapter 1 </strong> </p>
<p> Hello World1 </p>
<p> Hello World2 </p>
<p> Hello World3 </p>
<p> <strong> Chapter 2 </strong> </p>
<p> Hello World4 </p>
<p> Hello World5 </p>
<p> Hello World6 </p>
<p> <strong> Chapter 3 </strong> </p>
<p> Hello World7 </p>
<p> Hello World8 </p>
<p> Hello World9 </p>
</div>
'''

soup = BeautifulSoup(raw_html, 'html.parser')
content_root = soup.find('div', id='book')
chapter_list = []
temp_buffer = []

for p_tag in content_root.find_all('p'):
    # 识别章节标题:p标签内嵌套strong标签
    if p_tag.find('strong'):
        if temp_buffer:
            chapter_list.append('\n'.join(temp_buffer))
            temp_buffer = []
        continue
    # 普通段落直接保留原始HTML格式存入缓存
    temp_buffer.append(str(p_tag))
# 追加最后一个章节的内容
if temp_buffer:
    chapter_list.append('\n'.join(temp_buffer))

# 生成对应参数变量
param_1 = chapter_list[0]
param_2 = chapter_list[1]
param_3 = chapter_list[2]

# 输出验证结果
print("param_1 = \n")
print(param_1)
print("\nparam_2 = \n")
print(param_2)
print("\nparam_3 = \n")
print(param_3)

运行后输出的参数内容完全符合你要求的格式。

内容的提问来源于stack exchange,提问作者Jang

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.03 12:15:02