You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python爬取无按钮ID的动态网站?附国泰航空影视站实例

国泰航空娱乐影视页面爬取解决方案

问题背景

需爬取国泰航空影视页面的完整内容:https://entertainment.cathaypacific.com/catalog?template=movie&parent=%E9%9B%BB%E5%BD%B1。该页面采用动态加载机制,大部分内容需点击底部"Show More"按钮逐步加载,初始代码仅能获取少量数据,尝试Selenium点击按钮但未完成完整流程。


方案一:Selenium 循环加载完整内容

通过循环点击"Show More"按钮直至内容加载完毕,配合显式等待避免元素超时问题,同时优化数据提取逻辑:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.common.exceptions import NoSuchElementException
import pandas as pd

# 初始化Chrome浏览器
driver = webdriver.Chrome()
driver.get("https://entertainment.cathaypacific.com/catalog?template=movie&parent=%E9%9B%BB%E5%BD%B1")
driver.maximize_window()

# 循环点击Show More,直到按钮消失
while True:
    try:
        # 等待按钮可点击
        more_button = WebDriverWait(driver, 10).until(
            EC.element_to_be_clickable((By.CLASS_NAME, "showMore"))
        )
        more_button.click()
        # 等待新内容加载完成
        WebDriverWait(driver, 10).until(
            EC.staleness_of(more_button)
        )
    except NoSuchElementException:
        # 无更多内容,终止循环
        break

# 提取影视卡片数据
movie_cards = driver.find_elements(By.CLASS_NAME, "catalog-item")
data_list = []
for card in movie_cards:
    title = card.find_element(By.TAG_NAME, "a").text.strip()
    info_elements = card.find_elements(By.CLASS_NAME, "ng-star-inserted")
    info1 = info_elements[0].text.strip() if len(info_elements) > 0 else ""
    info2 = info_elements[1].text.strip() if len(info_elements) > 1 else ""
    data_list.append({"影片标题": title, "信息1": info1, "信息2": info2})

# 保存为Excel文件
df = pd.DataFrame(data_list)
df.to_excel("国泰影视全列表.xlsx", index=False)

# 关闭浏览器
driver.quit()

方案二:Pyppeteer 自动加载优化(原工具适配)

基于你最初使用的Pyppeteer实现自动点击加载,无需切换工具链:

import asyncio
from pyppeteer import launch
from bs4 import BeautifulSoup
import pandas as pd

async def crawl_full_movies():
    browser = await launch(headless=False)
    page = await browser.newPage()
    await page.goto("https://entertainment.cathaypacific.com/catalog?template=movie&parent=%E9%9B%BB%E5%BD%B1", waitUntil="networkidle0")

    # 循环点击Show More加载全部内容
    while True:
        try:
            button = await page.querySelector(".showMore")
            if not button:
                break
            await page.click(".showMore")
            # 等待页面加载完成
            await page.waitForNavigation(waitUntil="networkidle0", timeout=10000)
        except Exception:
            break

    # 解析完整页面内容
    html = await page.content()
    soup = BeautifulSoup(html, 'html.parser')

    # 提取影视数据
    data_list = []
    movie_cards = soup.find_all(class_="catalog-item")
    for card in movie_cards:
        title = card.find("a").get_text(strip=True)
        info_elements = card.find_all(class_="ng-star-inserted")
        info1 = info_elements[0].get_text(strip=True) if info_elements else ""
        info2 = info_elements[1].get_text(strip=True) if len(info_elements) >=2 else ""
        data_list.append({"影片标题": title, "信息1": info1, "信息2": info2})

    # 保存数据
    df = pd.DataFrame(data_list)
    df.to_excel("国泰影视全列表.xlsx", index=False)
    
    await browser.close()

asyncio.get_event_loop().run_until_complete(crawl_full_movies())

核心优化点

  • 显式等待机制:避免页面未加载完成时操作元素,降低异常概率
  • 循环终止条件:通过检查按钮是否存在判断是否还有更多内容,防止无限循环
  • 数据提取逻辑:直接定位影视卡片容器(.catalog-item),相比原代码的列表切片更稳定,不易受页面结构变化影响

内容的提问来源于stack exchange,提问作者Hi Hi try

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.29 20:47:28