如何抓取网站隐藏内容?解决国泰航空娱乐页Show More加载问题
问题描述
目标抓取国泰航空娱乐网站(https://entertainment.cathaypacific.com/catalog?template=movie&parent=%E9%9B%BB%E5%BD%B1)的全部电影数据,需获取字段:
- 电影标题
- 年份
- 时长
- 对应href链接
现有问题:页面底部的「Show More」按钮限制了单次获取的数据量,当前代码只能拿到有限条目。
现有代码:
import asyncio from pyppeteer import launch from bs4 import BeautifulSoup import pandas as pd data1=[] data2=[] async def get_data(): browser = await launch(headless=False) page = await browser.newPage() await page.goto("https://entertainment.cathaypacific.com/catalog?template=movie&parent=%E9%9B%BB%E5%BD%B1", waitUntil="networkidle0") html = await page.content() soup = BeautifulSoup(html, 'html.parser') titles = soup.find_all('a') for title in titles[7:]: data1.append(title.text) infs = soup.find_all("span",{"class":"ng-star-inserted"}) for inf in infs: data2.append(inf.text.strip()) asyncio.get_event_loop().run_until_complete(get_data()) data1_e = data1[0:-5][::2] data1_o = data1[1:-5][::2] data2_e = data2[0:-5][::2] data2_o = data2[1:-5][::2] d1 = pd.DataFrame(data1_o) d2 = pd.DataFrame(data2_e) d3 = pd.DataFrame(data2_o) result = pd.concat([d1,d2,d3], axis=1,join='inner') print(result) df = pd.DataFrame(result).to_excel('HIHI.xlsx', index = False)
解决方案与代码优化建议
一、无需点击「Show More」获取全量数据
该网站基于Angular构建,数据通过异步API加载。直接抓取后台数据接口即可绕过前端分页限制:
- 打开浏览器开发者工具(F12),切换到「Network」标签,筛选「XHR」类型请求
- 点击「Show More」时,会发现类似
https://entertainment.cathaypacific.com/api/v2/content/catalog的请求,该接口支持分页参数(page、size) - 设置足够大的
size参数(比如size=1000),即可一次性获取全量电影数据
优化后的API直调代码
import requests import pandas as pd def fetch_all_movies(): url = "https://entertainment.cathaypacific.com/api/v2/content/catalog" params = { "parent": "电影", "template": "movie", "page": 0, "size": 1000 # 根据实际总数据量调整,可先调小测试 } headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36" } response = requests.get(url, params=params, headers=headers) response.raise_for_status() data = response.json() movie_list = [] for item in data["items"]: movie_info = { "标题": item["title"], "年份": item.get("year"), "时长": item.get("duration"), "href": f"https://entertainment.cathaypacific.com/content/{item['contentId']}" } movie_list.append(movie_info) return movie_list if __name__ == "__main__": movies = fetch_all_movies() df = pd.DataFrame(movies) print(df) df.to_excel("国泰航空电影数据.xlsx", index=False)
二、原代码的优化建议
如果坚持使用Pyppeteer+BeautifulSoup的方式,可从以下几点优化:
- 自动加载全量数据:循环点击「Show More」直到按钮消失,确保所有数据加载完成后再解析
- 精准定位元素:避免依赖索引切片提取数据(如
titles[7:]),改用元素的父容器或关联属性匹配,提升代码稳定性 - 合理化数据结构:用字典存储单条电影的完整信息,避免分散到多个列表后再拼接
- 资源自动清理:使用
async with管理浏览器和页面,确保资源自动释放
优化后的Pyppeteer版本代码
import asyncio from pyppeteer import launch from bs4 import BeautifulSoup import pandas as pd async def fetch_all_movies(): async with launch(headless=True) as browser: async with browser.newPage() as page: await page.setUserAgent("Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36") await page.goto("https://entertainment.cathaypacific.com/catalog?template=movie&parent=%E9%9B%BB%E5%BD%B1", waitUntil="networkidle0") # 自动点击Show More直到按钮消失 while True: try: btn = await page.querySelector("button.show-more-btn") if not btn: break await page.click("button.show-more-btn") await page.waitForNavigation(waitUntil="networkidle0") except: break html = await page.content() soup = BeautifulSoup(html, 'html.parser') movie_items = soup.find_all("div", class_="col-md-4 col-lg-3 ng-star-inserted") movie_list = [] for item in movie_items: title_tag = item.find("a", class_="title ng-star-inserted") info_spans = item.find_all("span", class_="ng-star-inserted") movie_info = { "标题": title_tag.text.strip() if title_tag else "", "年份": info_spans[0].text.strip() if len(info_spans) > 0 else "", "时长": info_spans[1].text.strip() if len(info_spans) > 1 else "", "href": f"https://entertainment.cathaypacific.com{title_tag['href']}" if title_tag else "" } movie_list.append(movie_info) return movie_list if __name__ == "__main__": movies = asyncio.get_event_loop().run_until_complete(fetch_all_movies()) df = pd.DataFrame(movies) print(df) df.to_excel("国泰航空电影数据.xlsx", index=False)
内容的提问来源于stack exchange,提问作者Hi Hi try
相关产品推荐
相关产品推荐

