You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

网站存在启动页(Splash Screen)时如何抓取URL链接?

问题说明

原Python爬虫代码可从https://ascscotties.com/抓取含roster的有效链接,但访问https://auyellowjackets.com/时,因目标网站存在**启动页(Splash Screen)**导致抓取失效。核心原因是启动页依赖JavaScript渲染或需要交互跳转,而requests库仅能获取静态HTML,无法处理动态加载的主站内容。

解决方法

使用支持JavaScript渲染的浏览器自动化工具(如Selenium),模拟真实浏览器加载页面,获取渲染后的主站HTML后再进行链接解析。

步骤1:安装依赖

先安装Selenium库及对应浏览器驱动(以Chrome为例):

pip install selenium

同时下载与本地Chrome版本匹配的ChromeDriver,确保其路径可被Python调用(或配置到系统环境变量中)。

步骤2:修改后的爬虫代码

from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from bs4 import BeautifulSoup
import re
import requests

R = []
url = "https://auyellowjackets.com/"
headers = {'User-Agent': 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10.6; rv:16.0) Gecko/20100101 Firefox/16.0'}

# 配置Chrome无头模式(无界面运行)
chrome_options = Options()
chrome_options.add_argument("--headless=new")
chrome_options.add_argument("--disable-gpu")
chrome_options.add_argument(f"user-agent={headers['User-Agent']}")

# 启动浏览器加载页面
driver = webdriver.Chrome(options=chrome_options)
driver.get(url)
# 等待页面加载完成(可根据实际情况调整等待时长)
driver.implicitly_wait(10)

# 获取渲染后的页面HTML并关闭浏览器
page_source = driver.page_source
driver.quit()

# 后续链接解析逻辑与原代码一致
soup = BeautifulSoup(page_source, 'html.parser')
links = soup.find_all('a', href=re.compile("roster"))
# 处理相对链接,拼接为完整URL
s = [
    link.get("href") if link.get("href").startswith("http") 
    else url.rstrip("/") + link.get("href") 
    for link in links
]

# 验证链接有效性并收集结果
for i in s:
    r = requests.get(i, allow_redirects=True, headers=headers)
    if r.status_code < 400:
        R.append(r.url)

# 输出结果
print(R)

补充说明

  • 若不想使用Selenium,也可通过抓包工具分析启动页的跳转逻辑,直接获取主站的真实URL,但这种方法稳定性差,网站更新跳转规则后会失效。
  • Selenium支持Firefox、Edge等多种浏览器,只需更换对应的驱动和配置即可适配。

内容的提问来源于stack exchange,提问作者vijish madhavan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.23 15:18:05