You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python+Selenium通过类名定位元素返回空字符串问题

解决Selenium爬取Surfline冲浪高度返回空字符串的问题

问题原因

你的目标元素结构是<span class="quiver-surf-height">3-4<sup>FT</sup></span>,用.text返回空大概率是两个原因:要么元素还没完全渲染就被你读取了,要么Selenium的.text读取逻辑和页面动态渲染机制不兼容。

可行解决方案

1. 先等元素完全加载再读取

Surfline是动态渲染的网站,必须加显式等待确保元素可见:

from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

# 补全浏览器初始化
options = Options()
driver = webdriver.Chrome(options=options)
driver.get("你的Surfline目标页面URL")

# 最长等10秒,直到元素可见
wait = WebDriverWait(driver, 10)
surf_element = wait.until(EC.visibility_of_element_located((By.CLASS_NAME, "quiver-surf-height")))

# 再获取文本
surf = surf_element.text
print(f"The contents of surf is: {surf}")

2. 换用DOM属性读取文本

如果.text还是空,直接调用页面DOM的文本属性试试:

# 用textContent获取所有子节点的文本内容
surf = surf_element.get_attribute("textContent").strip()
# 或者用innerText,效果类似
# surf = surf_element.get_attribute("innerText").strip()
print(f"The contents of surf is: {surf}")

3. 单独提取数字部分

如果只需要3-4而不需要FT,可以直接定位span的直接文本节点:

# 通过JS脚本获取直接文本内容
surf = driver.execute_script("return document.querySelector('.quiver-surf-height').firstChild.textContent;")
print(f"The contents of surf is: {surf}")

额外提醒

  • 确认你的浏览器已经正确打开目标页面,要是遇到Cloudflare验证这类反爬,得先过验证才能继续爬取。
  • 如果元素是被CSS隐藏的(比如display:none),visibility_of_element_located会自动等待元素变为可见,避免读取隐藏元素的空文本。

内容的提问来源于stack exchange,提问作者Jujimufoo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.17 09:10:35