You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python Selenium从雅虎财经AAPL新闻页筛选提取有效新闻文章链接

筛选抓取苹果相关新闻链接的实现方法

核心思路

不要全页抓取所有<a>标签,先通过页面结构特征定位新闻专属容器,再结合URL规则二次过滤,彻底排除广告和无关跳转链接。

具体操作步骤

  • 第一步:定位新闻列表专属区域
    雅虎财经个股新闻页的所有有效新闻项,都统一嵌套在class属性包含news-stream的容器下,每个新闻项的链接都有统一的路径前缀特征。
  • 第二步:添加URL规则过滤
    所有有效的单篇新闻链接,都符合https://finance.yahoo.com/news/前缀的特征,广告、导航类链接均不满足该规则。
  • 第三步:优化页面内容加载逻辑
    该页面为滚动加载模式,若需要抓取更多历史新闻,可以模拟页面向下滚动触发内容加载。

优化后可运行代码

from selenium import webdriver
from selenium.webdriver.common.by import By
import time
import requests
from bs4 import BeautifulSoup

# 初始化驱动
driver = webdriver.Chrome(executable_path='C:\\Users\\Home\\OneDrive\\Desktop\\AJ\\chromedriver_win32\\chromedriver.exe')
driver.get("https://finance.yahoo.com/quote/AAPL/news?p=AAPL")
# 可选:模拟滚动加载更多新闻,滚动3次可以拿到约30条新闻
for _ in range(3):
    driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
    time.sleep(2)

links = []
# 仅抓取新闻容器下的a标签
news_a_tags = driver.find_elements(By.XPATH, '//div[contains(@class,"news-stream")]//a')
for a in news_a_tags:
    href = a.get_attribute('href')
    # 过滤符合新闻前缀的链接,同时去重
    if href.startswith("https://finance.yahoo.com/news/") and href not in links:
        links.append(href)

# 请求头配置,避免被反爬拦截
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
}

def get_info(url):
    response = requests.get(url, headers=headers)
    if response.status_code != 200:
        return None, None, None
    soup = BeautifulSoup(response.text, 'lxml')
    try:
        news = soup.find('div', attrs={'class': 'caas-body'}).text
        headline = soup.find('h1').text 
        date = soup.find('time').text
        return news, headline, date
    except AttributeError:
        return None, None, None

# 测试抓取
for link in links[:5]:
    news, headline, date = get_info(link)
    if headline:
        print(f"标题:{headline}\n日期:{date}\n内容摘要:{news[:100]}...\n---")

driver.quit()

注意事项

  • 若请求返回403错误,可以将requests请求替换为用driver.get(url)直接加载页面,再从driver中获取页面源码解析,规避反爬限制。
  • 页面结构可能随平台更新调整,若定位失效可以重新审查元素,更新容器的class特征即可。

内容的提问来源于stack exchange,提问作者huy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.04 04:48:01