如何爬取纯HTML长文章中按strong标签分章节的段落并存储
实现方案
你可以直接用Python的BeautifulSoup库完成提取,操作步骤如下:
- 先安装依赖库
pip install beautifulsoup4
- 运行以下代码即可输出你需要的变量内容:
from bs4 import BeautifulSoup # 替换为你自己的原始HTML长文本 raw_html = ''' <div id="book"> <h3>This is book</h3> <p> <strong> Chapter 1 </strong> </p> <p> Hello World1 </p> <p> Hello World2 </p> <p> Hello World3 </p> <p> <strong> Chapter 2 </strong> </p> <p> Hello World4 </p> <p> Hello World5 </p> <p> Hello World6 </p> <p> <strong> Chapter 3 </strong> </p> <p> Hello World7 </p> <p> Hello World8 </p> <p> Hello World9 </p> </div> ''' soup = BeautifulSoup(raw_html, 'html.parser') content_root = soup.find('div', id='book') chapter_list = [] temp_buffer = [] for p_tag in content_root.find_all('p'): # 识别章节标题:p标签内嵌套strong标签 if p_tag.find('strong'): if temp_buffer: chapter_list.append('\n'.join(temp_buffer)) temp_buffer = [] continue # 普通段落直接保留原始HTML格式存入缓存 temp_buffer.append(str(p_tag)) # 追加最后一个章节的内容 if temp_buffer: chapter_list.append('\n'.join(temp_buffer)) # 生成对应参数变量 param_1 = chapter_list[0] param_2 = chapter_list[1] param_3 = chapter_list[2] # 输出验证结果 print("param_1 = \n") print(param_1) print("\nparam_2 = \n") print(param_2) print("\nparam_3 = \n") print(param_3)
运行后输出的参数内容完全符合你要求的格式。
内容的提问来源于stack exchange,提问作者Jang
相关产品推荐
相关产品推荐

