如何利用Wayback Machine爬取USA Today新闻标题与链接
问题
需要从Wayback Machine存档的2022年1月1日USA Today页面中爬取新闻标题及对应URL。
目标片段特征:
<a class="section-helper-flex section-helper-row ten-column spacer-small p1-container" data-index="2" data-module-name="promo-story-thumb-small" href="https://www.usatoday.com/story/sports/nfl/2021/12/31/nfl-predictions-wrong-worst-preseason-mvp-super-bowl/9061437002/" onclick="firePromoAnalytics(event)"><div class="section-helper-flex section-helper-column p1-text-wrap"><div class="promo-premium-content-label-wrap"></div><div class="p1-title"><div class="p1-title-spacer">Revisiting our worst NFL predictions of 2021: What went wrong?</div></div><div class="p1-info-wrap"><span class="p1-label">NFL</span>
尝试的代码(无法精准提取):
import re from bs4 import BeautifulSoup from selenium import webdriver from selenium.webdriver.chrome.options import Options import requests url = 'https://web.archive.org/web/20220101001435/https://www.usatoday.com/' options = Options() options.headless = False driver = webdriver.Chrome(options=options) driver.get(url) soup = BeautifulSoup(driver.page_source, "html.parser") links=soup.find_all('a',href=True) print (links)
期望输出为二维数组:
[["Revisiting our worst NFL predictions of 2021: What went wrong?","https://www.usatoday.com/story/sports/nfl/2021/12/31/nfl-predictions-wrong-worst-preseason-mvp-super-bowl/9061437002/"], ...]
解决方案
核心思路
- 精准定位目标a标签:筛选带有
p1-container类的a标签(目标新闻项的外层容器) - 对每个符合条件的a标签,提取
href属性作为新闻链接 - 在a标签内部找到
p1-title-spacer类元素,提取其文本作为新闻标题 - 将标题和链接组合为子列表,最终收集成二维数组
修改后的代码
import re from bs4 import BeautifulSoup from selenium import webdriver from selenium.webdriver.chrome.options import Options import requests url = 'https://web.archive.org/web/20220101001435/https://www.usatoday.com/' options = Options() options.headless = False driver = webdriver.Chrome(options=options) driver.get(url) soup = BeautifulSoup(driver.page_source, "html.parser") # 筛选目标a标签:仅抓取含p1-container类的元素 target_links = soup.find_all('a', class_='p1-container', href=True) result = [] for link in target_links: # 提取内部标题元素 title_elem = link.find('div', class_='p1-title-spacer') # 确保标题和链接都存在再添加结果 if title_elem and link['href']: title = title_elem.get_text(strip=True) news_url = link['href'] result.append([title, news_url]) # 输出最终二维数组结果 print(result) # 关闭浏览器进程 driver.quit()
补充说明
- 用
class_='p1-container'直接过滤出新闻项容器,避免抓取导航栏、广告等无关链接 get_text(strip=True)会自动去除标题文本中的多余空格、换行符,让结果更整洁- 增加判断条件防止页面结构异常导致的报错,提升代码稳定性
内容的提问来源于stack exchange,提问作者wintermelon
相关产品推荐
相关产品推荐

