You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Selenium爬取US News高中排名页无法加载全量100条数据怎么办

解决US News全美最佳高中排名前100条数据抓取问题

现有代码问题排查

  • 按钮定位错误:你使用find_element_by_id匹配的pager__ButtonContentContainer-sc-1i8e93j-3 zIUhv是元素的class属性值,而非id,定位逻辑完全错误,无法触发加载更多操作
  • 未处理反爬检测:默认启动的ChromeDriver会被US News的反爬策略识别为自动化工具,限制返回数据条数
  • 滚动逻辑冗余:页面不需要滚动到指定位置也可以触发加载更多按钮的点击操作

修复后的Selenium实现方案

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from bs4 import BeautifulSoup
import time

# 配置Chrome参数,隐藏自动化特征
options = webdriver.ChromeOptions()
options.add_argument("--disable-blink-features=AutomationControlled")
options.add_experimental_option("excludeSwitches", ["enable-automation"])
options.add_experimental_option("useAutomationExtension", False)

url = "https://www.usnews.com/education/best-high-schools/search?national-rank-range-min=1&national-rank-range-max=100"
driver = webdriver.Chrome(options=options)
# 修改webdriver指纹,避免被识别
driver.execute_cdp_cmd("Page.addScriptToEvaluateOnNewDocument", {
    "source": "Object.defineProperty(navigator, 'webdriver', {get: () => undefined})"
})

driver.get(url)
time.sleep(5)

# 总共需要点击加载更多4次,每次加载20条,凑满100条
for _ in range(4):
    try:
        # 正确定位加载更多按钮,用CSS选择器匹配class
        loadmore = WebDriverWait(driver, 10).until(
            EC.element_to_be_clickable((By.CSS_SELECTOR, "button[class*='pager__Button']"))
        )
        loadmore.click()
        time.sleep(2) # 等待新数据加载完成
    except Exception as e:
        print("没有更多可加载内容:", e)
        break

html_source = driver.page_source
soup = BeautifulSoup(html_source,'html.parser')
driver.quit()

# 后续自行从soup中提取排名数据即可

更稳定的接口调用方案

这个方案不需要模拟浏览器操作,直接请求US News的后端数据接口,参数设置返回条数为100即可一次性拿到全部数据:

import requests

headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36",
    "Accept": "application/json, text/plain, */*"
}
# 接口参数可自定义返回条数,这里设置size为100直接返回前100条
params = {
    "national-rank-range-min": 1,
    "national-rank-range-max": 100,
    "page": 1,
    "size": 100
}
response = requests.get("https://www.usnews.com/education/best-high-schools/search/api/rankings", headers=headers, params=params)
data = response.json()
# 排名数据存储在data的items字段中,直接提取即可
rank_list = data.get("items", [])

这个方案效率更高,也不会出现页面渲染导致的数据缺失问题。

内容的提问来源于stack exchange,提问作者Nal

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.29 07:06:04