You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Selenium爬取纽约大都会歌剧院演出日程页面问题求助

解决方案:爬取大都会歌剧院动态生成的演出日程

Hey there! 作为刚接触爬虫的新手,遇到JS渲染页面的提取问题太正常了,别着急,咱们一步步来解决你的问题——核心就是确保页面完全加载+找对元素的定位方式,结合Selenium和BeautifulSoup的优势就能拿到你要的剧目名称、日期和时间啦。

第一步:先搞懂问题出在哪

你说只能拿到作曲家信息,大概率是这两个原因:

  • 页面还没完全渲染就开始提取内容,JS生成的核心元素还没加载出来
  • 用了错误的选择器(比如选到了静态加载的作曲家元素,而动态生成的剧目/日期元素用了不同的class/id)

第二步:具体实现步骤

1. 配置Selenium,确保页面完全加载

别用生硬的time.sleep(),改用显式等待让Selenium等到目标元素出现再操作;如果页面是滚动加载更多演出,还要模拟滚动到底部加载全部内容。

2. 用开发者工具定位正确的元素

打开大都会的日程页面,按F12打开开发者工具,切换到Elements标签,找到单个演出项的HTML结构——比如你会看到每个演出项都有一个父容器,里面包含剧目名称、日期时间的子元素,记下它们的class或id(比如performance-title、performance-datetime这类)。

3. 完整代码示例

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from bs4 import BeautifulSoup
import time

# 配置Chrome浏览器,避免被反爬检测
options = webdriver.ChromeOptions()
options.add_experimental_option("excludeSwitches", ["enable-automation"])
options.add_experimental_option('useAutomationExtension', False)
# 可选:开启无头模式,让浏览器在后台运行
# options.add_argument('--headless=new')

# 初始化浏览器并访问目标页面
driver = webdriver.Chrome(options=options)
target_url = "这里替换成大都会歌剧院的日程页面URL"
driver.get(target_url)

# 显式等待:直到演出列表的父容器加载完成(最多等10秒)
wait = WebDriverWait(driver, 10)
try:
    performance_container = wait.until(
        EC.presence_of_element_located((By.CLASS_NAME, "这里替换成演出列表的class名"))
    )
    print("页面核心内容加载完成!")
except Exception as e:
    print(f"等待加载失败:{e}")
    driver.quit()

# 模拟滚动,加载所有演出内容(如果页面是滚动加载的话)
last_scroll_height = driver.execute_script("return document.body.scrollHeight")
while True:
    # 滚动到底部
    driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
    # 等待2秒让新内容加载
    time.sleep(2)
    new_scroll_height = driver.execute_script("return document.body.scrollHeight")
    # 如果高度不再变化,说明所有内容加载完毕
    if new_scroll_height == last_scroll_height:
        break
    last_scroll_height = new_scroll_height

# 获取渲染后的页面源码,交给BeautifulSoup解析
page_html = driver.page_source
soup = BeautifulSoup(page_html, "html.parser")

# 提取核心信息
performances = []
for item in soup.find_all(class_="这里替换成单个演出项的class名"):
    # 提取剧目名称(替换成你找到的剧目元素class)
    title = item.find(class_="performance-title").get_text(strip=True) if item.find(class_="performance-title") else "未知剧目"
    # 提取日期时间(替换成你找到的日期时间元素class)
    date_time = item.find(class_="performance-datetime").get_text(strip=True) if item.find(class_="performance-datetime") else "未知时间"
    
    performances.append({
        "剧目名称": title,
        "演出日期时间": date_time
    })

# 打印结果或者导出到日历格式(比如ICS)
for perf in performances:
    print(f"【剧目】{perf['剧目名称']} | 【时间】{perf['演出日期时间']}")

# 关闭浏览器
driver.quit()

关键注意事项

  • 替换class名:代码里的这里替换成XXX部分,一定要用你在开发者工具里找到的真实class/id,比如大都会的演出项可能用event-card、title-link这类命名
  • ChromeDriver版本匹配:要确保你的ChromeDriver版本和本地Chrome浏览器版本一致,不然会启动失败
  • 反爬应对:如果遇到访问被限制,可以尝试添加options.add_argument('--user-agent="Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36..."')模拟真实浏览器请求

如果按照上面的步骤还是拿不到元素,你可以把开发者工具里单个演出项的HTML结构贴出来,这样能更精准地帮你调整选择器~

内容的提问来源于stack exchange,提问作者Jack Fleeting

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 07:21:37