使用BeautifulSoup抓取247sports排名历史链接遇错,求解决方法
问题:提取247sports页面中的排名历史链接
你当前代码的问题在于:Soup.find_all('ul', {'class':'ranks-list'})返回的是BeautifulSoup元素的列表,而列表对象没有find_all()方法,所以直接调用past_link.find_all('a')会报错。
解决方案
方法1:先获取列表中的单个ul元素,再查找目标链接
从返回的列表中取出唯一的ul元素(从你的输出看页面只有一个该类名的ul),再在这个ul下查找带rank-history-link类的a标签:
import requests from bs4 import BeautifulSoup # 补充定义headers,避免被网站拦截 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' } url = 'https://247sports.com/Player/Trevor-Lawrence-61350/college-212444/' pageTree = requests.get(url, headers=headers) Soup = BeautifulSoup(pageTree.content, 'html.parser') # find_all返回列表,取第一个ul元素 past_link_list = Soup.find_all('ul', {'class':'ranks-list'}) if past_link_list: past_link = past_link_list[0] # 定位带指定类名的a标签 history_link = past_link.find('a', class_='rank-history-link') if history_link: target_url = history_link['href'] print(target_url)
方法2:直接定位目标a标签(更高效)
跳过查找ul的步骤,直接通过类名定位目标a标签,一步到位:
import requests from bs4 import BeautifulSoup headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' } url = 'https://247sports.com/Player/Trevor-Lawrence-61350/college-212444/' pageTree = requests.get(url, headers=headers) Soup = BeautifulSoup(pageTree.content, 'html.parser') # 直接查找带rank-history-link类的a标签 history_link = Soup.find('a', class_='rank-history-link') if history_link: target_url = history_link['href'] print(target_url)
关键说明
find_all()返回的是元素列表,必须通过索引(如[0])取出单个元素后,才能调用元素的find()/find_all()方法- 使用
class_参数指定类名(因为class是Python关键字,不能直接作为参数名) - 补充
headers是为了模拟浏览器请求,避免被网站反爬机制拦截
内容的提问来源于stack exchange,提问作者BLuta
相关产品推荐
相关产品推荐

