You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Databricks中爬取Yahoo Finance加拿大站TECK-B.TO的2024年全部新闻

问题分析

Yahoo Finance的新闻列表采用动态加载机制,初始HTTP请求仅返回页面顶部的少量新闻(即你当前爬取到的2条),剩余新闻需在用户滚动页面时,由前端向服务器请求补充数据。直接使用requests+BeautifulSoup只能获取初始HTML中的静态内容,无法捕获动态加载的部分。


解决方案1:用Selenium模拟滚动加载所有新闻

步骤1:安装依赖

在Databricks Notebook中执行以下命令安装所需库:

%pip install selenium beautifulsoup4 pandas webdriver-manager

步骤2:修改后的完整代码

import time
import pandas as pd
from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from bs4 import BeautifulSoup
from pyspark.sql.types import StructType, StructField, StringType, TimestampType
from datetime import datetime

def fetch_yahoo_news_2024():
    url = "https://ca.finance.yahoo.com/quote/TECK-B.TO/news/"
    
    # 配置Chrome无头模式,适配Databricks无界面环境
    chrome_options = Options()
    chrome_options.add_argument("--headless=new")
    chrome_options.add_argument("--no-sandbox")
    chrome_options.add_argument("--disable-dev-shm-usage")
    chrome_options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36")
    
    # 初始化浏览器
    driver = webdriver.Chrome(options=chrome_options)
    driver.get(url)
    
    # 循环滚动加载所有新闻,直到页面高度不再变化
    last_height = driver.execute_script("return document.body.scrollHeight")
    while True:
        driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
        time.sleep(3)  # 等待新内容加载
        new_height = driver.execute_script("return document.body.scrollHeight")
        if new_height == last_height:
            break
        last_height = new_height
    
    # 解析完整页面源码
    soup = BeautifulSoup(driver.page_source, "html.parser")
    driver.quit()
    
    articles = []
    for item in soup.find_all('li', class_='js-stream-content'):
        link = item.find('a')['href'] if item.find('a') else None
        title = item.find('a').text.strip() if item.find('a') else None
        date_str = item.find('time')['datetime'] if item.find('time') else None
        
        if date_str:
            try:
                date = datetime.strptime(date_str, "%Y-%m-%dT%H:%M:%SZ")
                if date.year == 2024:
                    articles.append({
                        'title': title,
                        'link': f"https://ca.finance.yahoo.com{link}" if link else None,
                        'date': date
                    })
            except Exception as e:
                print(f"日期解析失败: {e}, 原始字符串: {date_str}")
                continue
    
    # 转换为Spark DataFrame
    schema = StructType([
        StructField("title", StringType(), nullable=True),
        StructField("link", StringType(), nullable=True),
        StructField("date", TimestampType(), nullable=True)
    ])
    
    if articles:
        return spark.createDataFrame(pd.DataFrame(articles), schema=schema)
    else:
        print("未找到2024年相关新闻")
        return spark.createDataFrame(pd.DataFrame(columns=['title', 'link', 'date']), schema=schema)

# 执行函数并显示结果
news_df = fetch_yahoo_news_2024()
display(news_df)

代码关键说明

  • 无头浏览器配置:--headless=new启用无界面模式,适配Databricks集群环境;--no-sandbox和--disable-dev-shm-usage解决Linux系统下的权限与内存限制问题。
  • 滚动加载逻辑:通过循环滚动页面到底部,等待新内容加载,直到页面高度不再变化,确保所有新闻被完整加载。
  • 容错处理:添加日期解析异常捕获,避免个别格式异常的日期导致程序中断。

解决方案2:直接调用Yahoo新闻API(可选)

Yahoo Finance的新闻加载依赖内部API接口(通常包含/v1/finance/news路径),可通过浏览器抓包获取请求参数后,直接用requests请求API获取全量数据。但需注意:

  • API请求需携带特定Cookie或Authorization头,容易过期失效。
  • 高频请求可能触发反爬限制,需控制请求频率。

注意事项

  • 反爬限制:Yahoo Finance有反爬机制,建议添加合理延迟(如time.sleep()),避免短时间内大量请求导致IP被封禁。
  • Databricks集群适配:若Selenium无法运行,需确保集群已安装ChromeDriver,或使用Databricks官方提供的浏览器驱动服务。

内容的提问来源于stack exchange,提问作者Learning-Developer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.18 00:14:53