使用Selenium/SeleniumBase时目标span标签为空无属性的问题求助
问题描述
我用Python爬取MLB.com的棒球新秀数据,手动打开浏览器看页面源码时,能看到目标JSON数据存在于一个<span>标签的data-init-state属性里;但用Selenium或SeleniumBase工具时,这个<span>标签始终是空的(就是<span></span>),完全没data-init-state属性,根本拿不到数据。
试过这些操作都没用:
- 用Selenium的
find_element()和SeleniumBase的get_element()找带data-init-state的span,一直报元素未找到 - 常规Selenium WebDriver,无头、有头模式都试了,把sleep延长到120秒,保存的页面源码里span还是空的
- 用SeleniumBase的
uc=True(防检测模式)+ 隐身窗口,同样sleep很久,结果依旧 - 有头模式下,脚本sleep的时候我手动右键看页面源码,span是有
data-init-state的,但脚本执行完保存的HTML里还是空的
怀疑网站检测到了自动化工具,阻止了填充这个span的脚本运行,求解决办法?
解决方案
1. 强化浏览器防检测(针对Selenium/SeleniumBase)
网站大概率是靠检测浏览器的自动化特征(比如navigator.webdriver标识、Chrome DevTools协议痕迹)来限制的,试试这些调整:
Selenium版本优化代码:
from selenium import webdriver from selenium.webdriver.chrome.options import Options from time import sleep url = "https://www.mlb.com/prospects" chrome_options = Options() # 要是需要无头模式就保留,不需要就删掉 chrome_options.add_argument('--headless=new') # 核心防检测参数 chrome_options.add_argument('--disable-blink-features=AutomationControlled') chrome_options.add_experimental_option("excludeSwitches", ["enable-automation"]) chrome_options.add_experimental_option('useAutomationExtension', False) # 模拟真实用户代理,也可以换成你自己浏览器的UA chrome_options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36") driver = webdriver.Chrome(options=chrome_options) # 执行JS移除webdriver标识,这步很关键 driver.execute_script("Object.defineProperty(navigator, 'webdriver', {get: () => undefined})") driver.get(url) sleep(10) prospects_source_html = driver.page_source # 记得加utf-8编码,避免乱码 with open('prospects_source_html.html', 'w', encoding='utf-8') as file: file.write(prospects_source_html) driver.quit()
SeleniumBase版本优化代码:
from seleniumbase import SB url = "https://www.mlb.com/prospects" # 配置更贴近真实浏览器的参数 with SB(uc=True, headed=True, incognito=True, user_agent="Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36") as sb: sb.uc_open_with_reconnect(url, 8) # 额外注入隐藏webdriver的JS sb.execute_script("Object.defineProperty(navigator, 'webdriver', {get: () => undefined})") sb.sleep(10) sb.save_page_source('prospects_source_html.html')
2. 直接抓API接口(更省心高效)
既然数据是JS动态填充的,肯定是从后台API拿的,没必要绕浏览器。步骤如下:
- 打开浏览器F12,切到Network面板,过滤
XHR/Fetch请求 - 刷新页面,找带
prospects或者player关键词的请求,复制它的URL和请求头 - 用
requests直接调用接口,示例代码:
import requests # 替换成你找到的真实API URL api_url = "https://xxxx.mlb.com/prospects/xxxx" headers = { "user-agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36", "referer": "https://www.mlb.com/prospects" } response = requests.get(api_url, headers=headers) data = response.json() # 直接处理JSON数据就行 print(data)
这种方法完全不会触发反爬,速度还快,优先推荐。
3. 智能等待元素加载(替代固定sleep)
要是非得用浏览器自动化,别用固定sleep,改成等待目标属性出现:
from selenium import webdriver from selenium.webdriver.chrome.options import Options from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC url = "https://www.mlb.com/prospects" chrome_options = Options() # 同样加防检测参数 chrome_options.add_argument('--disable-blink-features=AutomationControlled') chrome_options.add_experimental_option("excludeSwitches", ["enable-automation"]) chrome_options.add_experimental_option('useAutomationExtension', False) driver = webdriver.Chrome(options=chrome_options) driver.execute_script("Object.defineProperty(navigator, 'webdriver', {get: () => undefined})") driver.get(url) # 最多等30秒,直到span出现data-init-state属性 wait = WebDriverWait(driver, 30) target_span = wait.until( EC.presence_of_element_located((By.XPATH, "//span[@data-init-state]")) ) # 直接获取属性值 init_state_json = target_span.get_attribute("data-init-state") print(init_state_json) driver.quit()
内容的提问来源于stack exchange,提问作者bdr
相关产品推荐
相关产品推荐

