Kaggle页面<li>标签元素爬取失败,返回空列表求助
问题描述
尝试爬取Kaggle代码页面(https://www.kaggle.com/code?sortBy=voteCount&page=1)的信息,目标元素位于<ul role="list" class="km-list km-list--three-line">下的<li role="listitem" class="sc-jfmDQi hfJycS">标签中,想要提取每个元素的标题、点赞数、关联竞赛、评论数及链接,但执行代码后输出为空列表[]。
页面HTML示例
<ul role="list" class="km-list km-list--three-line"> <li role="listitem" class="sc-jfmDQi hfJycS"> <div class="sc-eKszNL ktNGam"> <div class="sc-hiMGwR GvHYb sc-czGAKf ffgKfy"> <div class="sc-ehmTmK bMKNkA"> <a href="/pmarcelino" target="_blank" class="sc-kgUAyh eqNroj" aria-label="Pedro Marcelino, PhD"> <div data-testid="avatar-image" title="Pedro Marcelino, PhD" class="sc-hTtwUo eLCpfL" style="background-image: url("https://storage.googleapis.com/kaggle-avatars/thumbnails/175415-gr.jpg");"></div> <svg width="64" height="64" viewBox="0 0 64 64"> <circle r="30.5" cx="32" cy="32" fill="none" stroke-width="3" style="stroke: rgb(241, 243, 244);"></circle> <path d="M 49.92745019492043 56.6750183284359 A 30.5 30.5 0 0 0 32 1.5" fill="none" stroke-width="3" style="stroke: rgb(32, 190, 255);"></path> </svg> </a> </div> </div> <a class="sc-lbOyJj eeGduD sc-jFAmCJ fNVSOc" href="/code/pmarcelino/comprehensive-data-exploration-with-python"> <div class="sc-ckMVTt jHrWZQ"> <div class="sc-iBkjds sc-fLlhyt sc-fbPSWO uVZhN izULIq A-dENW">Comprehensive data exploration with Python</div> <span class="sc-jIZahH sc-himrzO sc-fXynhf kdTVzc glCpMy ctwKCt"> <span><span>Updated <span title="Sat Apr 30 2022 21:20:37 GMT+0200 (heure d’été d’Europe centrale)" aria-label="7 months ago">7mo ago</span></span></span> </span> <span class="sc-jIZahH sc-himrzO sc-fXynhf kdTVzc glCpMy ctwKCt"> <span class="sc-bPPhlf hfaBPJ"> <a href="/code/pmarcelino/comprehensive-data-exploration-with-python/comments" class="sc-dPyBCJ sc-bBXxYQ sc-bOJcbE cSRCiy cFEurs gTFrUa">1819 comments</a> · <span class="sc-ibQCxQ idHgMS"> <span class="sc-cKajLJ jNrpDQ">House Prices - Advanced Regression Techniques</span> </span> </span> </span> </div> </a> <div class="sc-gFGZVQ jDMEwY sc-yTtWT kdALiC"> <div class="sc-dICTr dlQsbO"> <button mode="default" data-testid="upvotebutton__upvote" aria-label="Upvote" class="sc-dNezTh sc-lkcIho cSGKPD cTyEVx"> <i class="rmwc-icon rmwc-icon--ligature google-material-icons sc-gKXOVf jWACgA" sizevalue="18px">arrow_drop_up</i> </button> <span mode="default" class="sc-gXmSlM sc-cCsOjp sc-hAGLhy cKhlzA piYDj mWvOY">12770</span> </div> <span class="sc-dXxSUK cHoXAr"> <span class="sc-jIZahH sc-himrzO sc-hRwTwm kdTVzc glCpMy IISDK"> <img role="presentation" alt="" src="/static/images/medals/competitions/golds@1x.png" style="height: 9px; width: 9px;"> Gold </span> <div class="mdc-menu-surface--anchor"> <button aria-label="more_horiz" class="sc-jSMfEi eiMRSN sc-bgrGEg hPFZMI google-material-icons">more_horiz</button> </div> </span> </div> </div> <div class="sc-lbxAil LkNdN"></div> </li> </ul>
编写的Python代码
import pandas as pd import requests from bs4 import BeautifulSoup headers = {'User-Agent':'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_11_4) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/50.0.2661.94 Safari/537.36'} url = "https://www.kaggle.com/code?sortBy=voteCount&page=1" req = requests.get(url, headers = headers) soup = BeautifulSoup(req.text, 'html.parser') html_content = soup.find_all('li', attrs = {'class': 'sc-jfmDQi hfJycS'}) data = [] for elements in html_content: data.append({ 'title': elements.find("div", {"class": "sc-iBkjds sc-fLlhyt sc-fbPSWO uVZhN izULIq A-dENW"}).text, 'stars': elements.find("span", {"class": "sc-gXmSlM sc-cCsOjp sc-hAGLhy cKhlzA piYDj mWvOY"}).text, 'resume': elements.find("span", {"class": "sc-cKajLJ jNrpDQ"}).text, 'comments': elements.find("span", {"class": "sc-dPyBCJ sc-bBXxYQ sc-bOJcbE cSRCiy cFEurs gTFrUa"}).text, 'link': elements.get('href')}) print(data)
输出结果
[]
问题原因与解决办法
核心问题
- 动态内容渲染:Kaggle的代码列表是通过JavaScript动态加载的,
requests只能获取页面的静态HTML,无法拿到JS渲染后的实际内容,所以找不到目标元素。 - 动态类名不可靠:代码中使用的
sc-jfmDQi hfJycS这类类名是前端框架动态生成的,会随时变化,依赖这类类名定位元素会导致失效。 - 链接提取错误:目标链接不在
<li>标签上,而是在<li>内部的<a>标签中。
解决方案
方案一:使用Selenium获取动态渲染内容
Selenium可以模拟浏览器加载页面,获取完整的渲染后内容,适合处理动态页面:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from bs4 import BeautifulSoup url = "https://www.kaggle.com/code?sortBy=voteCount&page=1" # 初始化Chrome浏览器(需提前下载对应版本的ChromeDriver) driver = webdriver.Chrome() driver.get(url) # 等待目标列表加载完成,最多等待10秒 wait = WebDriverWait(driver, 10) wait.until(EC.presence_of_all_elements_located((By.CSS_SELECTOR, 'ul[role="list"].km-list--three-line li[role="listitem"]'))) # 获取渲染后的页面源码 soup = BeautifulSoup(driver.page_source, 'html.parser') driver.quit() data = [] # 基于稳定的属性和结构定位元素,避免依赖动态类名 items = soup.select('ul[role="list"].km-list--three-line li[role="listitem"]') for item in items: # 提取标题 title_elem = item.select_one('a[href^="/code/"] div:first-of-type') title = title_elem.text.strip() if title_elem else None # 提取点赞数 stars_elem = item.select_one('button[data-testid="upvotebutton__upvote"] + span') stars = stars_elem.text.strip() if stars_elem else None # 提取关联竞赛 resume_elem = item.select_one('span[class*="cKajLJ"]') resume = resume_elem.text.strip() if resume_elem else None # 提取评论数 comments_elem = item.select_one('a[href$="/comments"]') comments = comments_elem.text.strip() if comments_elem else None # 提取完整链接 link_elem = item.select_one('a[href^="/code/"]') link = f"https://www.kaggle.com{link_elem['href']}" if link_elem else None if title: data.append({ 'title': title, 'stars': stars, 'resume': resume, 'comments': comments, 'link': link }) print(data)
方案二:使用Kaggle API(推荐)
Kaggle提供了官方API,可以合规、稳定地获取代码数据,避免反爬问题:
- 安装Kaggle库:
pip install kaggle - 在Kaggle账号中创建API令牌,保存到本地指定路径
- 通过API获取代码列表,示例代码如下:
import kaggle # 初始化API(需提前配置好Kaggle令牌) kaggle.api.authenticate() # 获取按点赞排序的代码列表 codes = kaggle.api.kernels_list(sort_by='voteCount', page=1) data = [] for code in codes: data.append({ 'title': code.title, 'stars': code.voteCount, 'resume': code.competitionTitle if hasattr(code, 'competitionTitle') else None, 'comments': code.commentCount, 'link': f"https://www.kaggle.com/code/{code.ownerSlug}/{code.slug}" }) print(data)
注意事项
- 使用Selenium时,需确保浏览器驱动版本与浏览器版本匹配,且注意控制请求频率,避免触发Kaggle的反爬机制。
- 优先选择官方API方案,不仅稳定,还能避免违反Kaggle的使用条款。
内容的提问来源于stack exchange,提问作者ladybug
相关产品推荐
相关产品推荐

