如何用Python & Beautiful Soup提取特定标签类组合间的文章内容?
问题描述
作为网页爬取及Beautiful Soup新手,我需要从单页博客网站构建包含标题、内容、发布日期的DataFrame。目前已成功获取标题和发布日期,但文章内容提取有误——每个段落被识别为单独条目。
页面结构如下:
<h2 class = "thisYear" title = "Click here to display/hide information"> "First Post Title" </h2> <p class ="pubdate" style="display: block;"> 2022-07-11</p> <p style="display: block;"> "First paragraph of post"</p> <p style="display: block;"> "Second paragraph of post"</p> <h2 class = "thisYear" title = "Click here to display/hide information"> "Second Post Title" </h2> <p class ="pubdate" style="display: block;"> 2022-07-07</p> <p style="display: block;"> "First paragraph of post"</p> <p style="display: block;"> "Second paragraph of post"</p>
当前代码:
r = requests.get(URL,allow_redirects=True) soup = BeautifulSoup(r.content, 'html5lib') tag = 'p' title_class_name = "thisYear" news_class_name = "thisYear" date_class_name = "pubdate" df = pd.DataFrame() title_list = [] news_list =[] date_list = [] title_table = soup.findAll('h2',attrs= {'class':title_class_name}) news_table = soup.findAll(tag,attrs= {'class': None}) date_table = soup.findAll(tag,attrs= {'class':date_class_name}) for (title , news, date) in zip(title_table, news_table, date_table): title_list.append(title.text) news_list.append(news.text) date_list.append(date.text) df['title'] = title_list df['news']=news_list df['publish_date']=date_list df
核心问题:如何提取每个<h2 class="thisYear">标签之间的文章内容?
解决方案
问题出在你直接把所有无class的p标签和标题、日期做zip匹配,导致每个段落对应一个标题,而非将同一篇文章的所有段落合并。正确思路是遍历每个标题,收集该标题之后到下一个标题之前的所有内容段落,同时提取对应发布日期。
修改后的代码如下:
import requests from bs4 import BeautifulSoup import pandas as pd r = requests.get(URL, allow_redirects=True) soup = BeautifulSoup(r.content, 'html5lib') title_class_name = "thisYear" date_class_name = "pubdate" df = pd.DataFrame() title_list = [] news_list = [] date_list = [] # 获取所有标题节点 titles = soup.find_all('h2', class_=title_class_name) for title in titles: # 提取并清理标题文本 clean_title = title.text.strip() title_list.append(clean_title) # 定位标题后的发布日期节点 date_tag = title.find_next_sibling('p', class_=date_class_name) clean_date = date_tag.text.strip() date_list.append(clean_date) # 收集当前文章的所有内容段落 content_paragraphs = [] next_sibling = date_tag.find_next_sibling() # 循环遍历兄弟节点,直到遇到下一个h2标签停止 while next_sibling is not None and next_sibling.name != 'h2': if next_sibling.name == 'p' and not next_sibling.has_attr('class'): clean_p = next_sibling.text.strip() content_paragraphs.append(clean_p) next_sibling = next_sibling.find_next_sibling() # 将多个段落合并为单个字符串 full_content = '\n'.join(content_paragraphs) news_list.append(full_content) # 构建最终DataFrame df['title'] = title_list df['news'] = news_list df['publish_date'] = date_list print(df)
关键处理逻辑:
- 遍历单个标题节点,而非批量抓取所有p标签
- 使用
find_next_sibling精准定位发布日期和后续内容节点 - 通过循环收集内容,直到遇到下一个h2标签终止
- 将同一篇文章的所有段落合并为单个字符串,避免拆分多个条目
内容的提问来源于stack exchange,提问作者Dave
相关产品推荐
相关产品推荐

