使用BeautifulSoup抓取带Load More按钮的动态UFC选手页面遇阻求助
UFC官网选手数据抓取问题:无法获取“Load More”加载的内容
我尝试抓取UFC官网(https://www.ufc.com/athletes/all)的选手数据,但无法获取点击“Load More”按钮后加载的内容,使用BeautifulSoup库未能成功,恳请提供解决建议。
以下是我目前的代码:
from cgitb import html from selenium.webdriver.common.keys import Keys from bs4 import BeautifulSoup import requests html_text = requests.get('https://www.ufc.com/athletes/all').text soup = BeautifulSoup(html_text, "lxml") fighters = soup.find_all('div', class_ = ("node node--type-athlete node--view-mode-all-athletes-result ds-1col clearfix")) for fighter in fighters: fighter_name = fighter.find('span', class_ = ("c-listing-athlete__name")).text.replace("_", " ").replace("-", " ") fighter_nickname = fighter.find('div', class_ = ("field field--name-nickname field--type-string field--label-hidden")) fighter_nickname = fighter_nickname.text if fighter_nickname else None fighter_weight_class = fighter.find('div', class_ = ("field field--name-stats-weight-class field--type-entity-reference field--label-hidden field__items")).text fighter_ufc_record = fighter.find('span', class_ = ("c-listing-athlete__record")).text print(f'''Fighter Name: {fighter_name.strip()}''') print("************************")
问题原因
requests.get只能获取页面初始的静态HTML,而“Load More”加载的内容是通过AJAX动态请求获取的,BeautifulSoup无法解析这些未包含在初始HTML里的动态内容,所以只能拿到第一页的选手数据。
解决建议
方法1:用Selenium模拟浏览器点击加载
你已经导入了Selenium,直接用它模拟浏览器操作,循环点击“Load More”按钮直到所有内容加载完成,再解析页面:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from bs4 import BeautifulSoup import time # 初始化Chrome浏览器(需提前安装对应版本的chromedriver) driver = webdriver.Chrome() driver.get('https://www.ufc.com/athletes/all') # 循环点击Load More按钮,直到按钮消失 while True: try: # 等待按钮可点击后执行点击 load_more_btn = WebDriverWait(driver, 10).until( EC.element_to_be_clickable((By.XPATH, "//button[contains(text(), 'Load More')]")) ) load_more_btn.click() time.sleep(2) # 等待内容加载完成 except: # 按钮不存在时退出循环 break # 获取加载完成后的完整页面源码 page_source = driver.page_source driver.quit() # 用BeautifulSoup解析页面 soup = BeautifulSoup(page_source, "lxml") fighters = soup.find_all('div', class_ = ("node node--type-athlete node--view-mode-all-athletes-result ds-1col clearfix")) # 后续数据提取逻辑和原代码一致 for fighter in fighters: fighter_name = fighter.find('span', class_ = ("c-listing-athlete__name")).text.replace("_", " ").replace("-", " ") fighter_nickname = fighter.find('div', class_ = ("field field--name-nickname field--type-string field--label-hidden")) fighter_nickname = fighter_nickname.text if fighter_nickname else None fighter_weight_class = fighter.find('div', class_ = ("field field--name-stats-weight-class field--type-entity-reference field--label-hidden field__items")).text fighter_ufc_record = fighter.find('span', class_ = ("c-listing-athlete__record")).text print(f'''Fighter Name: {fighter_name.strip()}''') print("************************")
方法2:直接调用AJAX接口(效率更高)
打开浏览器开发者工具(F12)切换到“网络”标签,点击“Load More”时观察发送的请求,会发现页面通过API接口获取选手数据。直接用requests请求该接口获取JSON数据,无需模拟浏览器:
import requests url = "https://www.ufc.com/api/v3/athletes" params = { "page": 0, "limit": 24, "sort": "last_name", "filter[status]": "Active", "filter[weight_class]": "" } all_fighters = [] while True: response = requests.get(url, params=params) data = response.json() if not data: break all_fighters.extend(data) params["page"] += 1 # 解析JSON数据提取所需信息 for fighter in all_fighters: fighter_name = f"{fighter['first_name']} {fighter['last_name']}" fighter_nickname = fighter.get('nickname') fighter_weight_class = fighter.get('weight_class', {}).get('name') fighter_ufc_record = f"{fighter['wins']}-{fighter['losses']}-{fighter['draws']}" print(f'''Fighter Name: {fighter_name.strip()}''') print(f'''Nickname: {fighter_nickname}''') print(f'''Weight Class: {fighter_weight_class}''') print(f'''UFC Record: {fighter_ufc_record}''') print("************************")
注意:API接口参数可能会随网站更新变化,需自行抓包确认最新接口地址和参数。
内容的提问来源于stack exchange,提问作者Nick_Lopez43
相关产品推荐
相关产品推荐

