You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python & Beautiful Soup提取特定标签类组合间的文章内容?

问题描述

作为网页爬取及Beautiful Soup新手,我需要从单页博客网站构建包含标题、内容、发布日期的DataFrame。目前已成功获取标题和发布日期,但文章内容提取有误——每个段落被识别为单独条目。

页面结构如下:

<h2 class = "thisYear" title = "Click here to display/hide information">
"First Post Title" </h2>
<p class ="pubdate" style="display: block;"> 2022-07-11</p>
<p style="display: block;"> "First paragraph of post"</p>
<p style="display: block;"> "Second paragraph of post"</p>
<h2 class = "thisYear" title = "Click here to display/hide information">
"Second Post Title" </h2>
<p class ="pubdate" style="display: block;"> 2022-07-07</p>
<p style="display: block;"> "First paragraph of post"</p>
<p style="display: block;"> "Second paragraph of post"</p>

当前代码:

r = requests.get(URL,allow_redirects=True)
soup = BeautifulSoup(r.content, 'html5lib')
    
tag = 'p'
title_class_name = "thisYear"
news_class_name = "thisYear"
date_class_name = "pubdate"


df = pd.DataFrame()
title_list = []
news_list =[]
date_list = []

title_table = soup.findAll('h2',attrs= {'class':title_class_name})
news_table = soup.findAll(tag,attrs= {'class': None})
date_table = soup.findAll(tag,attrs= {'class':date_class_name})

for (title , news, date) in zip(title_table, news_table, date_table):
    title_list.append(title.text)
    news_list.append(news.text)
    date_list.append(date.text)
df['title'] = title_list
df['news']=news_list
df['publish_date']=date_list
df

核心问题:如何提取每个<h2 class="thisYear">标签之间的文章内容?

解决方案

问题出在你直接把所有无class的p标签和标题、日期做zip匹配,导致每个段落对应一个标题,而非将同一篇文章的所有段落合并。正确思路是遍历每个标题,收集该标题之后到下一个标题之前的所有内容段落,同时提取对应发布日期。

修改后的代码如下:

import requests
from bs4 import BeautifulSoup
import pandas as pd

r = requests.get(URL, allow_redirects=True)
soup = BeautifulSoup(r.content, 'html5lib')

title_class_name = "thisYear"
date_class_name = "pubdate"

df = pd.DataFrame()
title_list = []
news_list = []
date_list = []

# 获取所有标题节点
titles = soup.find_all('h2', class_=title_class_name)

for title in titles:
    # 提取并清理标题文本
    clean_title = title.text.strip()
    title_list.append(clean_title)
    
    # 定位标题后的发布日期节点
    date_tag = title.find_next_sibling('p', class_=date_class_name)
    clean_date = date_tag.text.strip()
    date_list.append(clean_date)
    
    # 收集当前文章的所有内容段落
    content_paragraphs = []
    next_sibling = date_tag.find_next_sibling()
    
    # 循环遍历兄弟节点,直到遇到下一个h2标签停止
    while next_sibling is not None and next_sibling.name != 'h2':
        if next_sibling.name == 'p' and not next_sibling.has_attr('class'):
            clean_p = next_sibling.text.strip()
            content_paragraphs.append(clean_p)
        next_sibling = next_sibling.find_next_sibling()
    
    # 将多个段落合并为单个字符串
    full_content = '\n'.join(content_paragraphs)
    news_list.append(full_content)

# 构建最终DataFrame
df['title'] = title_list
df['news'] = news_list
df['publish_date'] = date_list
print(df)

关键处理逻辑:

  • 遍历单个标题节点,而非批量抓取所有p标签
  • 使用find_next_sibling精准定位发布日期和后续内容节点
  • 通过循环收集内容,直到遇到下一个h2标签终止
  • 将同一篇文章的所有段落合并为单个字符串,避免拆分多个条目

内容的提问来源于stack exchange,提问作者Dave

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.26 06:18:29