You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

网页抓取求助:如何获取指定class下的电影标题、年份及时长?

解决国泰航空娱乐网站电影数据抓取问题

为什么你的代码抓不到数据?

这个网站采用了前端动态渲染技术(类似Angular框架),你用requests获取的只是页面的初始骨架模板,实际的电影数据是页面加载完成后通过JavaScript异步加载的,所以BeautifulSoup根本无法读取到这些动态生成的内容。

解决方案一:用Selenium模拟浏览器加载

Selenium可以模拟真实浏览器的加载流程,等页面完全渲染完成后再抓取内容,代码示例如下:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from bs4 import BeautifulSoup
import time

# 初始化Chrome浏览器(需提前安装对应版本的chromedriver)
driver = webdriver.Chrome()
driver.get('https://entertainment.cathaypacific.com/catalog?template=movie&parent=%E9%9B%BB%E5%BD%B1&language=yue')

try:
    # 等待电影元素加载出现,超时时间10秒
    WebDriverWait(driver, 10).until(
        EC.presence_of_element_located((By.CLASS_NAME, "ng-star-inserted"))
    )
    time.sleep(2)  # 额外等待2秒确保所有数据加载完毕

    # 获取渲染后的页面源码并解析
    soup = BeautifulSoup(driver.page_source, 'html.parser')
    movie_items = soup.find_all(class_="ng-star-inserted")

    # 格式化输出表头
    print(f"{'Title':<30} {'Year':<8} {'Length':<8}")
    # 遍历提取每部电影的信息
    for item in movie_items:
        title = item.find(class_="title").text.strip() if item.find(class_="title") else "N/A"
        year = item.find(class_="year").text.strip() if item.find(class_="year") else "N/A"
        length = item.find(class_="duration").text.strip() if item.find(class_="duration") else "N/A"
        print(f"{title:<30} {year:<8} {length:<8}")

finally:
    driver.quit()  # 执行完毕关闭浏览器

解决方案二:直接调用API接口(更高效)

通过浏览器开发者工具的「网络」面板,可以找到网站加载电影数据的API接口,直接请求这个接口获取JSON格式的数据,无需处理前端渲染逻辑:

import requests

# 替换为实际找到的API接口(需自行在浏览器网络请求中确认)
api_url = "https://entertainment.cathaypacific.com/api/catalog/movies?language=yue"
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"
}

response = requests.get(api_url, headers=headers)
if response.status_code == 200:
    movie_data = response.json()
    # 格式化输出表头
    print(f"{'Title':<30} {'Year':<8} {'Length':<8}")
    # 遍历JSON数据提取信息
    for movie in movie_data.get("items", []):
        title = movie.get("title", "N/A")
        year = movie.get("releaseYear", "N/A")
        length = movie.get("duration", "N/A")
        print(f"{title:<30} {year:<8} {length:<8}")
else:
    print(f"请求API失败,状态码:{response.status_code}")

注:API接口可能会随网站更新变化,需自行在浏览器的网络请求列表中查找带有api标识、返回JSON格式电影数据的请求地址。

内容的提问来源于stack exchange,提问作者Hi Hi try

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.29 17:32:57