网页抓取求助:如何获取指定class下的电影标题、年份及时长?
解决国泰航空娱乐网站电影数据抓取问题
为什么你的代码抓不到数据?
这个网站采用了前端动态渲染技术(类似Angular框架),你用requests获取的只是页面的初始骨架模板,实际的电影数据是页面加载完成后通过JavaScript异步加载的,所以BeautifulSoup根本无法读取到这些动态生成的内容。
解决方案一:用Selenium模拟浏览器加载
Selenium可以模拟真实浏览器的加载流程,等页面完全渲染完成后再抓取内容,代码示例如下:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from bs4 import BeautifulSoup import time # 初始化Chrome浏览器(需提前安装对应版本的chromedriver) driver = webdriver.Chrome() driver.get('https://entertainment.cathaypacific.com/catalog?template=movie&parent=%E9%9B%BB%E5%BD%B1&language=yue') try: # 等待电影元素加载出现,超时时间10秒 WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.CLASS_NAME, "ng-star-inserted")) ) time.sleep(2) # 额外等待2秒确保所有数据加载完毕 # 获取渲染后的页面源码并解析 soup = BeautifulSoup(driver.page_source, 'html.parser') movie_items = soup.find_all(class_="ng-star-inserted") # 格式化输出表头 print(f"{'Title':<30} {'Year':<8} {'Length':<8}") # 遍历提取每部电影的信息 for item in movie_items: title = item.find(class_="title").text.strip() if item.find(class_="title") else "N/A" year = item.find(class_="year").text.strip() if item.find(class_="year") else "N/A" length = item.find(class_="duration").text.strip() if item.find(class_="duration") else "N/A" print(f"{title:<30} {year:<8} {length:<8}") finally: driver.quit() # 执行完毕关闭浏览器
解决方案二:直接调用API接口(更高效)
通过浏览器开发者工具的「网络」面板,可以找到网站加载电影数据的API接口,直接请求这个接口获取JSON格式的数据,无需处理前端渲染逻辑:
import requests # 替换为实际找到的API接口(需自行在浏览器网络请求中确认) api_url = "https://entertainment.cathaypacific.com/api/catalog/movies?language=yue" headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36" } response = requests.get(api_url, headers=headers) if response.status_code == 200: movie_data = response.json() # 格式化输出表头 print(f"{'Title':<30} {'Year':<8} {'Length':<8}") # 遍历JSON数据提取信息 for movie in movie_data.get("items", []): title = movie.get("title", "N/A") year = movie.get("releaseYear", "N/A") length = movie.get("duration", "N/A") print(f"{title:<30} {year:<8} {length:<8}") else: print(f"请求API失败,状态码:{response.status_code}")
注:API接口可能会随网站更新变化,需自行在浏览器的网络请求列表中查找带有api标识、返回JSON格式电影数据的请求地址。
内容的提问来源于stack exchange,提问作者Hi Hi try
相关产品推荐
相关产品推荐

