使用Python Selenium从雅虎财经AAPL新闻页筛选提取有效新闻文章链接
筛选抓取苹果相关新闻链接的实现方法
核心思路
不要全页抓取所有<a>标签,先通过页面结构特征定位新闻专属容器,再结合URL规则二次过滤,彻底排除广告和无关跳转链接。
具体操作步骤
- 第一步:定位新闻列表专属区域
雅虎财经个股新闻页的所有有效新闻项,都统一嵌套在class属性包含news-stream的容器下,每个新闻项的链接都有统一的路径前缀特征。 - 第二步:添加URL规则过滤
所有有效的单篇新闻链接,都符合https://finance.yahoo.com/news/前缀的特征,广告、导航类链接均不满足该规则。 - 第三步:优化页面内容加载逻辑
该页面为滚动加载模式,若需要抓取更多历史新闻,可以模拟页面向下滚动触发内容加载。
优化后可运行代码
from selenium import webdriver from selenium.webdriver.common.by import By import time import requests from bs4 import BeautifulSoup # 初始化驱动 driver = webdriver.Chrome(executable_path='C:\\Users\\Home\\OneDrive\\Desktop\\AJ\\chromedriver_win32\\chromedriver.exe') driver.get("https://finance.yahoo.com/quote/AAPL/news?p=AAPL") # 可选:模拟滚动加载更多新闻,滚动3次可以拿到约30条新闻 for _ in range(3): driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") time.sleep(2) links = [] # 仅抓取新闻容器下的a标签 news_a_tags = driver.find_elements(By.XPATH, '//div[contains(@class,"news-stream")]//a') for a in news_a_tags: href = a.get_attribute('href') # 过滤符合新闻前缀的链接,同时去重 if href.startswith("https://finance.yahoo.com/news/") and href not in links: links.append(href) # 请求头配置,避免被反爬拦截 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } def get_info(url): response = requests.get(url, headers=headers) if response.status_code != 200: return None, None, None soup = BeautifulSoup(response.text, 'lxml') try: news = soup.find('div', attrs={'class': 'caas-body'}).text headline = soup.find('h1').text date = soup.find('time').text return news, headline, date except AttributeError: return None, None, None # 测试抓取 for link in links[:5]: news, headline, date = get_info(link) if headline: print(f"标题:{headline}\n日期:{date}\n内容摘要:{news[:100]}...\n---") driver.quit()
注意事项
- 若请求返回403错误,可以将
requests请求替换为用driver.get(url)直接加载页面,再从driver中获取页面源码解析,规避反爬限制。 - 页面结构可能随平台更新调整,若定位失效可以重新审查元素,更新容器的class特征即可。
内容的提问来源于stack exchange,提问作者huy
相关产品推荐
相关产品推荐

