代码无法提取输出Greek Rank网站的兄弟会/姐妹会名称,无报错
问题分析与修复方案
核心原因
Greek Rank网站会拦截无浏览器标识的请求,直接用requests.get()发送请求会返回不含目标内容的页面,导致cards为空;另外原代码依赖的页面结构可能已更新,选择器不匹配。
修复步骤
1. 添加请求头绕过反爬
给请求添加User-Agent模拟浏览器访问,确保获取到完整页面内容:
import requests from bs4 import BeautifulSoup URL = "https://www.greekrank.com/uni/51/greek-life/" # 模拟Chrome浏览器请求头 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } page = requests.get(URL, headers=headers) # 先验证请求是否成功,正常应返回200 print(page.status_code) soup = BeautifulSoup(page.content, 'html.parser')
2. 修正页面元素选择器
原代码中card-body类名可能已失效,需先打印soup.prettify()查看实际页面结构,替换为正确的选择器。例如当前页面中排名项的容器类名为rank-item,排名和名称的标签类名如下:
# 替换为实际页面中的元素容器类名 cards = soup.find_all('div', class_='rank-item') print("Top 5 Fraternities/Sororities:") print("-----------------------------------") print("| Ranking | Fraternity/Sorority |") print("-----------------------------------") for i, card in enumerate(cards): if i >= 5: break # 匹配实际页面的排名、名称标签类名 ranking = card.find('span', class_='rank').text.strip() name = card.find('h3', class_='org-name').text.strip() print(f"| {ranking:<8} | {name:<23} |") print("-----------------------------------")
3. 处理动态加载内容(若需)
如果页面内容是JavaScript动态渲染的,requests无法抓取,需用selenium模拟浏览器加载:
from selenium import webdriver from selenium.webdriver.chrome.service import Service from webdriver_manager.chrome import ChromeDriverManager from bs4 import BeautifulSoup # 初始化Chrome浏览器 driver = webdriver.Chrome(service=Service(ChromeDriverManager().install())) driver.get("https://www.greekrank.com/uni/51/greek-life/") # 等待页面加载完成 driver.implicitly_wait(10) # 获取渲染后的页面源码 soup = BeautifulSoup(driver.page_source, 'html.parser') cards = soup.find_all('div', class_='rank-item') # 后续打印逻辑同前 driver.quit()
排查技巧
- 先检查
page.status_code,确认请求是否成功(返回200) - 打印
soup.prettify()查看实际返回的HTML结构,验证选择器是否匹配 - 若页面有动态内容,必须用浏览器自动化工具抓取渲染后的页面
内容的提问来源于stack exchange,提问作者Alex
相关产品推荐
相关产品推荐

