如何在Databricks中爬取Yahoo Finance加拿大站TECK-B.TO的2024年全部新闻
问题分析
Yahoo Finance的新闻列表采用动态加载机制,初始HTTP请求仅返回页面顶部的少量新闻(即你当前爬取到的2条),剩余新闻需在用户滚动页面时,由前端向服务器请求补充数据。直接使用requests+BeautifulSoup只能获取初始HTML中的静态内容,无法捕获动态加载的部分。
解决方案1:用Selenium模拟滚动加载所有新闻
步骤1:安装依赖
在Databricks Notebook中执行以下命令安装所需库:
%pip install selenium beautifulsoup4 pandas webdriver-manager
步骤2:修改后的完整代码
import time import pandas as pd from selenium import webdriver from selenium.webdriver.chrome.options import Options from bs4 import BeautifulSoup from pyspark.sql.types import StructType, StructField, StringType, TimestampType from datetime import datetime def fetch_yahoo_news_2024(): url = "https://ca.finance.yahoo.com/quote/TECK-B.TO/news/" # 配置Chrome无头模式,适配Databricks无界面环境 chrome_options = Options() chrome_options.add_argument("--headless=new") chrome_options.add_argument("--no-sandbox") chrome_options.add_argument("--disable-dev-shm-usage") chrome_options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36") # 初始化浏览器 driver = webdriver.Chrome(options=chrome_options) driver.get(url) # 循环滚动加载所有新闻,直到页面高度不再变化 last_height = driver.execute_script("return document.body.scrollHeight") while True: driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") time.sleep(3) # 等待新内容加载 new_height = driver.execute_script("return document.body.scrollHeight") if new_height == last_height: break last_height = new_height # 解析完整页面源码 soup = BeautifulSoup(driver.page_source, "html.parser") driver.quit() articles = [] for item in soup.find_all('li', class_='js-stream-content'): link = item.find('a')['href'] if item.find('a') else None title = item.find('a').text.strip() if item.find('a') else None date_str = item.find('time')['datetime'] if item.find('time') else None if date_str: try: date = datetime.strptime(date_str, "%Y-%m-%dT%H:%M:%SZ") if date.year == 2024: articles.append({ 'title': title, 'link': f"https://ca.finance.yahoo.com{link}" if link else None, 'date': date }) except Exception as e: print(f"日期解析失败: {e}, 原始字符串: {date_str}") continue # 转换为Spark DataFrame schema = StructType([ StructField("title", StringType(), nullable=True), StructField("link", StringType(), nullable=True), StructField("date", TimestampType(), nullable=True) ]) if articles: return spark.createDataFrame(pd.DataFrame(articles), schema=schema) else: print("未找到2024年相关新闻") return spark.createDataFrame(pd.DataFrame(columns=['title', 'link', 'date']), schema=schema) # 执行函数并显示结果 news_df = fetch_yahoo_news_2024() display(news_df)
代码关键说明
- 无头浏览器配置:
--headless=new启用无界面模式,适配Databricks集群环境;--no-sandbox和--disable-dev-shm-usage解决Linux系统下的权限与内存限制问题。 - 滚动加载逻辑:通过循环滚动页面到底部,等待新内容加载,直到页面高度不再变化,确保所有新闻被完整加载。
- 容错处理:添加日期解析异常捕获,避免个别格式异常的日期导致程序中断。
解决方案2:直接调用Yahoo新闻API(可选)
Yahoo Finance的新闻加载依赖内部API接口(通常包含/v1/finance/news路径),可通过浏览器抓包获取请求参数后,直接用requests请求API获取全量数据。但需注意:
- API请求需携带特定Cookie或Authorization头,容易过期失效。
- 高频请求可能触发反爬限制,需控制请求频率。
注意事项
- 反爬限制:Yahoo Finance有反爬机制,建议添加合理延迟(如
time.sleep()),避免短时间内大量请求导致IP被封禁。 - Databricks集群适配:若Selenium无法运行,需确保集群已安装ChromeDriver,或使用Databricks官方提供的浏览器驱动服务。
内容的提问来源于stack exchange,提问作者Learning-Developer
相关产品推荐
相关产品推荐

