You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

卫报环境板块Wildlife栏目标题爬取问题求助(2022年10-12月)

《卫报》Wildlife板块爬虫问题排查与修复

一、列表页爬取无结果的原因及修复

问题根源

  1. 动态内容未加载:《卫报》主页的板块内容依赖JavaScript动态渲染,直接用requests.get获取的初始HTML里,gu-island标签的props属性是JSON格式字符串,而非可直接访问的字典,你之前的代码直接取element['props']['id']会报错或匹配不到内容。
  2. 板块入口错误:你访问的是环境频道主页,wildlife板块2022年10-12月的历史内容需要直接访问对应归档页,主页仅展示最新内容,不会加载过去三个月的归档数据。
  3. 标签选择器偏差:即使拿到动态内容,wildlife板块的文章标题也不是用h1标签,而是嵌套在h3或h2标签里,需要匹配特定类的容器。

修复后的列表页爬取代码

import requests
import json
from bs4 import BeautifulSoup
from datetime import datetime

# 遍历2022年10-12月的归档页
months = ['oct', 'nov', 'dec']
month_numbers = [10, 11, 12]
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'
}

for idx, month in enumerate(months):
    target_month = month_numbers[idx]
    url = f'https://www.theguardian.com/environment/wildlife/2022/{month}'
    r = requests.get(url, headers=headers)
    soup = BeautifulSoup(r.text, 'html.parser')
    
    # 解析gu-island的props属性
    islands = soup.find_all('gu-island')
    for island in islands:
        props_str = island.get('props', '')
        if not props_str:
            continue
        try:
            props = json.loads(props_str)
            # 筛选wildlife板块的内容
            if props.get('id') == 'environment/wildlife':
                # 提取文章容器
                article_containers = island.find_all('div', class_='dcr-12w641e')
                for container in article_containers:
                    # 获取标题和副标题
                    headline = container.find('h3').text.strip()
                    standfirst_elem = container.find('p', class_='dcr-1m7i1f9')
                    standfirst = standfirst_elem.text.strip() if standfirst_elem else '无副标题'
                    
                    # 校验文章日期是否在目标月份
                    date_str = container.find('time')['datetime']
                    article_date = datetime.fromisoformat(date_str.replace('Z', '+00:00'))
                    if article_date.year == 2022 and article_date.month == target_month:
                        print(f"标题: {headline}")
                        print(f"副标题: {standfirst}")
                        print("---")
        except json.JSONDecodeError:
            continue

二、单篇文章爬取冗余内容的原因及修复

问题根源

你用soup.find_all('p')会匹配页面中所有<p>标签,包括正文段落、页脚说明等冗余内容,而副标题(standfirst)有专属的类名,需要精准定位。

修复后的单篇文章爬取代码

import requests
from bs4 import BeautifulSoup

url = 'https://www.theguardian.com/environment/2022/dec/30/tales-of-killer-wild-boar-in-uk-are-hogwash-say-environmentalists'
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'
}
r = requests.get(url, headers=headers)
soup = BeautifulSoup(r.text, 'html.parser')

# 精准定位标题和副标题
headline = soup.find('h1', class_='dcr-1395u3b').text.strip()
standfirst = soup.find('p', class_='dcr-1m7i1f9').text.strip()

print(f"标题: {headline}")
print(f"副标题: {standfirst}")

内容的提问来源于stack exchange,提问作者arles

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 10:35:25