You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup的find_all返回空列表,无法获取冲浪网站浪高数据

Surfline浪高数据爬取问题解决

你的代码返回空列表主要有两个核心原因:

  1. 网页动态渲染:Surfline的浪高数据通过JavaScript动态加载,requests.get()仅能获取静态HTML源码,目标元素在初始响应中并不存在。
  2. 动态类名失效:你使用的mui-style-1sd06dx是Material UI生成的动态类名,每次请求都会变化,无法作为稳定的定位标识。

方案1:调用官方API(推荐)

Surfline的数据通过API接口返回,直接请求API比解析HTML更可靠。示例代码如下:

import requests

# 目标地点的浪高预测API(从浏览器开发者工具Network面板获取)
url = "https://services.surfline.com/kbyg/spots/forecasts/wave?spotId=5842041f4e65fad6a7708e84&days=5&intervalHours=1"
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"
}

response = requests.get(url, headers=headers)
data = response.json()

# 提取并打印浪高数据
for forecast in data['data']['wave']:
    timestamp = forecast['timestamp']
    min_height = forecast['surf']['min']
    max_height = forecast['surf']['max']
    print(f"时间: {timestamp}, 浪高范围: {min_height} - {max_height} 英尺")

方案2:用Selenium加载动态页面

若不想查找API,可通过Selenium模拟浏览器加载完整页面后解析:

from selenium import webdriver
from selenium.webdriver.common.by import By
from bs4 import BeautifulSoup
import time

# 初始化Chrome浏览器(需提前下载对应版本的ChromeDriver)
driver = webdriver.Chrome()
driver.get("https://www.surfline.com/surf-report/meia-praia/5842041f4e65fad6a7708e84?view=table")

# 等待页面动态内容加载完成
time.sleep(3)

# 获取完整页面源码并解析
page_source = driver.page_source
soup = BeautifulSoup(page_source, "html.parser")

# 通过稳定的类名定位浪高元素(避开动态生成的类)
height_elements = soup.find_all(class_="SurfCell_height__FkNCP")
for element in height_elements:
    print("浪高:", element.get_text())

# 关闭浏览器
driver.quit()

注意事项

  • API接口参数可能随网站更新变动,建议通过浏览器开发者工具(Network面板,过滤XHR请求)实时获取最新接口。
  • 使用Selenium时,需确保浏览器驱动版本与本地浏览器版本匹配,且调整等待时间以保证页面加载完整。

内容的提问来源于stack exchange,提问作者kiko bangarang

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.12 17:04:57