You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何利用Wayback Machine爬取USA Today新闻标题与链接

问题

需要从Wayback Machine存档的2022年1月1日USA Today页面中爬取新闻标题及对应URL。

目标片段特征:

<a class="section-helper-flex section-helper-row ten-column spacer-small p1-container" data-index="2" data-module-name="promo-story-thumb-small" href="https://www.usatoday.com/story/sports/nfl/2021/12/31/nfl-predictions-wrong-worst-preseason-mvp-super-bowl/9061437002/" onclick="firePromoAnalytics(event)"><div class="section-helper-flex section-helper-column p1-text-wrap"><div class="promo-premium-content-label-wrap"></div><div class="p1-title"><div class="p1-title-spacer">Revisiting our worst NFL predictions of 2021: What went wrong?</div></div><div class="p1-info-wrap"><span class="p1-label">NFL</span>

尝试的代码(无法精准提取):

import re
from bs4 import BeautifulSoup
from selenium import webdriver
from selenium.webdriver.chrome.options import Options
import requests

url = 'https://web.archive.org/web/20220101001435/https://www.usatoday.com/'
options = Options()
options.headless = False
driver = webdriver.Chrome(options=options)
driver.get(url)
soup = BeautifulSoup(driver.page_source, "html.parser")

links=soup.find_all('a',href=True)
print (links)

期望输出为二维数组:

[["Revisiting our worst NFL predictions of 2021: What went wrong?","https://www.usatoday.com/story/sports/nfl/2021/12/31/nfl-predictions-wrong-worst-preseason-mvp-super-bowl/9061437002/"], ...]

解决方案

核心思路

  1. 精准定位目标a标签:筛选带有p1-container类的a标签(目标新闻项的外层容器)
  2. 对每个符合条件的a标签,提取href属性作为新闻链接
  3. 在a标签内部找到p1-title-spacer类元素,提取其文本作为新闻标题
  4. 将标题和链接组合为子列表,最终收集成二维数组

修改后的代码

import re
from bs4 import BeautifulSoup
from selenium import webdriver
from selenium.webdriver.chrome.options import Options
import requests

url = 'https://web.archive.org/web/20220101001435/https://www.usatoday.com/'
options = Options()
options.headless = False
driver = webdriver.Chrome(options=options)
driver.get(url)
soup = BeautifulSoup(driver.page_source, "html.parser")

# 筛选目标a标签:仅抓取含p1-container类的元素
target_links = soup.find_all('a', class_='p1-container', href=True)
result = []

for link in target_links:
    # 提取内部标题元素
    title_elem = link.find('div', class_='p1-title-spacer')
    # 确保标题和链接都存在再添加结果
    if title_elem and link['href']:
        title = title_elem.get_text(strip=True)
        news_url = link['href']
        result.append([title, news_url])

# 输出最终二维数组结果
print(result)
# 关闭浏览器进程
driver.quit()

补充说明

  • 用class_='p1-container'直接过滤出新闻项容器,避免抓取导航栏、广告等无关链接
  • get_text(strip=True)会自动去除标题文本中的多余空格、换行符,让结果更整洁
  • 增加判断条件防止页面结构异常导致的报错,提升代码稳定性

内容的提问来源于stack exchange,提问作者wintermelon

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 10:22:49