使用Selenium爬取US News高中排名页无法加载全量100条数据怎么办
解决US News全美最佳高中排名前100条数据抓取问题
现有代码问题排查
- 按钮定位错误:你使用
find_element_by_id匹配的pager__ButtonContentContainer-sc-1i8e93j-3 zIUhv是元素的class属性值,而非id,定位逻辑完全错误,无法触发加载更多操作 - 未处理反爬检测:默认启动的ChromeDriver会被US News的反爬策略识别为自动化工具,限制返回数据条数
- 滚动逻辑冗余:页面不需要滚动到指定位置也可以触发加载更多按钮的点击操作
修复后的Selenium实现方案
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from bs4 import BeautifulSoup import time # 配置Chrome参数,隐藏自动化特征 options = webdriver.ChromeOptions() options.add_argument("--disable-blink-features=AutomationControlled") options.add_experimental_option("excludeSwitches", ["enable-automation"]) options.add_experimental_option("useAutomationExtension", False) url = "https://www.usnews.com/education/best-high-schools/search?national-rank-range-min=1&national-rank-range-max=100" driver = webdriver.Chrome(options=options) # 修改webdriver指纹,避免被识别 driver.execute_cdp_cmd("Page.addScriptToEvaluateOnNewDocument", { "source": "Object.defineProperty(navigator, 'webdriver', {get: () => undefined})" }) driver.get(url) time.sleep(5) # 总共需要点击加载更多4次,每次加载20条,凑满100条 for _ in range(4): try: # 正确定位加载更多按钮,用CSS选择器匹配class loadmore = WebDriverWait(driver, 10).until( EC.element_to_be_clickable((By.CSS_SELECTOR, "button[class*='pager__Button']")) ) loadmore.click() time.sleep(2) # 等待新数据加载完成 except Exception as e: print("没有更多可加载内容:", e) break html_source = driver.page_source soup = BeautifulSoup(html_source,'html.parser') driver.quit() # 后续自行从soup中提取排名数据即可
更稳定的接口调用方案
这个方案不需要模拟浏览器操作,直接请求US News的后端数据接口,参数设置返回条数为100即可一次性拿到全部数据:
import requests headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36", "Accept": "application/json, text/plain, */*" } # 接口参数可自定义返回条数,这里设置size为100直接返回前100条 params = { "national-rank-range-min": 1, "national-rank-range-max": 100, "page": 1, "size": 100 } response = requests.get("https://www.usnews.com/education/best-high-schools/search/api/rankings", headers=headers, params=params) data = response.json() # 排名数据存储在data的items字段中,直接提取即可 rank_list = data.get("items", [])
这个方案效率更高,也不会出现页面渲染导致的数据缺失问题。
内容的提问来源于stack exchange,提问作者Nal
相关产品推荐
相关产品推荐

