卫报环境板块Wildlife栏目标题爬取问题求助(2022年10-12月)
《卫报》Wildlife板块爬虫问题排查与修复
一、列表页爬取无结果的原因及修复
问题根源
- 动态内容未加载:《卫报》主页的板块内容依赖JavaScript动态渲染,直接用
requests.get获取的初始HTML里,gu-island标签的props属性是JSON格式字符串,而非可直接访问的字典,你之前的代码直接取element['props']['id']会报错或匹配不到内容。 - 板块入口错误:你访问的是环境频道主页,wildlife板块2022年10-12月的历史内容需要直接访问对应归档页,主页仅展示最新内容,不会加载过去三个月的归档数据。
- 标签选择器偏差:即使拿到动态内容,wildlife板块的文章标题也不是用
h1标签,而是嵌套在h3或h2标签里,需要匹配特定类的容器。
修复后的列表页爬取代码
import requests import json from bs4 import BeautifulSoup from datetime import datetime # 遍历2022年10-12月的归档页 months = ['oct', 'nov', 'dec'] month_numbers = [10, 11, 12] headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' } for idx, month in enumerate(months): target_month = month_numbers[idx] url = f'https://www.theguardian.com/environment/wildlife/2022/{month}' r = requests.get(url, headers=headers) soup = BeautifulSoup(r.text, 'html.parser') # 解析gu-island的props属性 islands = soup.find_all('gu-island') for island in islands: props_str = island.get('props', '') if not props_str: continue try: props = json.loads(props_str) # 筛选wildlife板块的内容 if props.get('id') == 'environment/wildlife': # 提取文章容器 article_containers = island.find_all('div', class_='dcr-12w641e') for container in article_containers: # 获取标题和副标题 headline = container.find('h3').text.strip() standfirst_elem = container.find('p', class_='dcr-1m7i1f9') standfirst = standfirst_elem.text.strip() if standfirst_elem else '无副标题' # 校验文章日期是否在目标月份 date_str = container.find('time')['datetime'] article_date = datetime.fromisoformat(date_str.replace('Z', '+00:00')) if article_date.year == 2022 and article_date.month == target_month: print(f"标题: {headline}") print(f"副标题: {standfirst}") print("---") except json.JSONDecodeError: continue
二、单篇文章爬取冗余内容的原因及修复
问题根源
你用soup.find_all('p')会匹配页面中所有<p>标签,包括正文段落、页脚说明等冗余内容,而副标题(standfirst)有专属的类名,需要精准定位。
修复后的单篇文章爬取代码
import requests from bs4 import BeautifulSoup url = 'https://www.theguardian.com/environment/2022/dec/30/tales-of-killer-wild-boar-in-uk-are-hogwash-say-environmentalists' headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' } r = requests.get(url, headers=headers) soup = BeautifulSoup(r.text, 'html.parser') # 精准定位标题和副标题 headline = soup.find('h1', class_='dcr-1395u3b').text.strip() standfirst = soup.find('p', class_='dcr-1m7i1f9').text.strip() print(f"标题: {headline}") print(f"副标题: {standfirst}")
内容的提问来源于stack exchange,提问作者arles
相关产品推荐
相关产品推荐

