You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

用Python爬取QS世界大学排名失败 求优化现有BeautifulSoup抓取代码

QS排名页面数据抓取优化方案

问题根源

QS2022世界大学排名页面的院校、排名、指标数据均为JavaScript动态渲染生成,直接通过requests发起的HTTP请求只能拿到未执行JS的初始静态源代码,不存在目标uni-link类的a标签,因此无法匹配到对应内容。

优化实现方案

推荐使用Selenium模拟浏览器加载完整页面后再解析数据,具体实现如下:

前置依赖安装

执行以下命令安装所需库:

pip install selenium webdriver-manager beautifulsoup4

优化后代码

from bs4 import BeautifulSoup
from selenium import webdriver
from selenium.webdriver.chrome.service import Service
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from webdriver_manager.chrome import ChromeDriverManager
import time

# 目标页面地址
url = "https://www.topuniversities.com/university-rankings/world-university-rankings/2022"

# 初始化浏览器(自动适配Chrome驱动)
options = webdriver.ChromeOptions()
# 不需要可视化界面可开启无头模式
# options.add_argument('--headless=new')
driver = webdriver.Chrome(service=Service(ChromeDriverManager().install()), options=options)

driver.get(url)
# 等待院校列表元素加载完成,最长等待10秒
WebDriverWait(driver, 10).until(
    EC.presence_of_element_located((By.CLASS_NAME, "uni-link"))
)

# 若需要获取学术声誉数据,先点击切换到Rankings Indicators标签
# indicator_tab = driver.find_element(By.XPATH, "//div[text()='Rankings Indicators']")
# indicator_tab.click()
# time.sleep(2) # 等待指标数据加载完成

# 拿渲染后的页面源码解析
html = driver.page_source
soup = BeautifulSoup(html, 'html.parser')

# 抓取院校名称
uni_links = soup.find_all('a', class_="uni-link")
for item in uni_links:
    print(item.text.strip())

# 关闭浏览器
driver.quit()

补充说明

  • 分页数据需要额外模拟点击下一页按钮,循环抓取所有页面内容
  • 学术声誉数据在切换标签后,对应类名可通过浏览器F12开发者工具定位匹配,和院校名称抓取逻辑一致

内容的提问来源于stack exchange,提问作者dolphin20

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.25 11:15:09