You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup抓取带Load More按钮的动态UFC选手页面遇阻求助

UFC官网选手数据抓取问题:无法获取“Load More”加载的内容

我尝试抓取UFC官网(https://www.ufc.com/athletes/all)的选手数据,但无法获取点击“Load More”按钮后加载的内容,使用BeautifulSoup库未能成功,恳请提供解决建议。

以下是我目前的代码:

from cgitb import html
from selenium.webdriver.common.keys import Keys
from bs4 import BeautifulSoup
import requests 

html_text = requests.get('https://www.ufc.com/athletes/all').text
soup = BeautifulSoup(html_text, "lxml")
fighters = soup.find_all('div', class_ = ("node node--type-athlete node--view-mode-all-athletes-result ds-1col clearfix"))

for fighter in fighters:
 fighter_name = fighter.find('span', class_ = ("c-listing-athlete__name")).text.replace("_", " ").replace("-", " ")
 fighter_nickname = fighter.find('div', class_ = ("field field--name-nickname field--type-string field--label-hidden"))
 fighter_nickname = fighter_nickname.text if fighter_nickname else None 
 fighter_weight_class = fighter.find('div', class_ = ("field field--name-stats-weight-class field--type-entity-reference field--label-hidden field__items")).text
 fighter_ufc_record = fighter.find('span', class_ = ("c-listing-athlete__record")).text
 print(f'''Fighter Name: {fighter_name.strip()}''')
 print("************************")

问题原因

requests.get只能获取页面初始的静态HTML,而“Load More”加载的内容是通过AJAX动态请求获取的,BeautifulSoup无法解析这些未包含在初始HTML里的动态内容,所以只能拿到第一页的选手数据。

解决建议

方法1:用Selenium模拟浏览器点击加载

你已经导入了Selenium,直接用它模拟浏览器操作,循环点击“Load More”按钮直到所有内容加载完成,再解析页面:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from bs4 import BeautifulSoup
import time

# 初始化Chrome浏览器(需提前安装对应版本的chromedriver)
driver = webdriver.Chrome()
driver.get('https://www.ufc.com/athletes/all')

# 循环点击Load More按钮,直到按钮消失
while True:
    try:
        # 等待按钮可点击后执行点击
        load_more_btn = WebDriverWait(driver, 10).until(
            EC.element_to_be_clickable((By.XPATH, "//button[contains(text(), 'Load More')]"))
        )
        load_more_btn.click()
        time.sleep(2)  # 等待内容加载完成
    except:
        # 按钮不存在时退出循环
        break

# 获取加载完成后的完整页面源码
page_source = driver.page_source
driver.quit()

# 用BeautifulSoup解析页面
soup = BeautifulSoup(page_source, "lxml")
fighters = soup.find_all('div', class_ = ("node node--type-athlete node--view-mode-all-athletes-result ds-1col clearfix"))

# 后续数据提取逻辑和原代码一致
for fighter in fighters:
    fighter_name = fighter.find('span', class_ = ("c-listing-athlete__name")).text.replace("_", " ").replace("-", " ")
    fighter_nickname = fighter.find('div', class_ = ("field field--name-nickname field--type-string field--label-hidden"))
    fighter_nickname = fighter_nickname.text if fighter_nickname else None 
    fighter_weight_class = fighter.find('div', class_ = ("field field--name-stats-weight-class field--type-entity-reference field--label-hidden field__items")).text
    fighter_ufc_record = fighter.find('span', class_ = ("c-listing-athlete__record")).text
    print(f'''Fighter Name: {fighter_name.strip()}''')
    print("************************")

方法2:直接调用AJAX接口(效率更高)

打开浏览器开发者工具(F12)切换到“网络”标签,点击“Load More”时观察发送的请求,会发现页面通过API接口获取选手数据。直接用requests请求该接口获取JSON数据,无需模拟浏览器:

import requests

url = "https://www.ufc.com/api/v3/athletes"
params = {
    "page": 0,
    "limit": 24,
    "sort": "last_name",
    "filter[status]": "Active",
    "filter[weight_class]": ""
}

all_fighters = []

while True:
    response = requests.get(url, params=params)
    data = response.json()
    
    if not data:
        break
    
    all_fighters.extend(data)
    params["page"] += 1

# 解析JSON数据提取所需信息
for fighter in all_fighters:
    fighter_name = f"{fighter['first_name']} {fighter['last_name']}"
    fighter_nickname = fighter.get('nickname')
    fighter_weight_class = fighter.get('weight_class', {}).get('name')
    fighter_ufc_record = f"{fighter['wins']}-{fighter['losses']}-{fighter['draws']}"
    print(f'''Fighter Name: {fighter_name.strip()}''')
    print(f'''Nickname: {fighter_nickname}''')
    print(f'''Weight Class: {fighter_weight_class}''')
    print(f'''UFC Record: {fighter_ufc_record}''')
    print("************************")

注意:API接口参数可能会随网站更新变化,需自行抓包确认最新接口地址和参数。

内容的提问来源于stack exchange,提问作者Nick_Lopez43

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.21 10:54:18