You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Kaggle页面<li>标签元素爬取失败,返回空列表求助

问题描述

尝试爬取Kaggle代码页面(https://www.kaggle.com/code?sortBy=voteCount&page=1)的信息,目标元素位于<ul role="list" class="km-list km-list--three-line">下的<li role="listitem" class="sc-jfmDQi hfJycS">标签中,想要提取每个元素的标题、点赞数、关联竞赛、评论数及链接,但执行代码后输出为空列表[]。

页面HTML示例

<ul role="list" class="km-list km-list--three-line">
  <li role="listitem" class="sc-jfmDQi hfJycS">
    <div class="sc-eKszNL ktNGam">
      <div class="sc-hiMGwR GvHYb sc-czGAKf ffgKfy">
        <div class="sc-ehmTmK bMKNkA">
          <a href="/pmarcelino" target="_blank" class="sc-kgUAyh eqNroj" aria-label="Pedro Marcelino, PhD">
            <div data-testid="avatar-image" title="Pedro Marcelino, PhD" class="sc-hTtwUo eLCpfL" style="background-image: url(&quot;https://storage.googleapis.com/kaggle-avatars/thumbnails/175415-gr.jpg&quot;);"></div>
            <svg width="64" height="64" viewBox="0 0 64 64">
              <circle r="30.5" cx="32" cy="32" fill="none" stroke-width="3" style="stroke: rgb(241, 243, 244);"></circle>
              <path d="M 49.92745019492043 56.6750183284359 A 30.5 30.5 0 0 0 32 1.5" fill="none" stroke-width="3" style="stroke: rgb(32, 190, 255);"></path>
            </svg>
          </a>
        </div>
      </div>
      <a class="sc-lbOyJj eeGduD sc-jFAmCJ fNVSOc" href="/code/pmarcelino/comprehensive-data-exploration-with-python">
        <div class="sc-ckMVTt jHrWZQ">
          <div class="sc-iBkjds sc-fLlhyt sc-fbPSWO uVZhN izULIq A-dENW">Comprehensive data exploration with Python</div>
          <span class="sc-jIZahH sc-himrzO sc-fXynhf kdTVzc glCpMy ctwKCt">
            <span><span>Updated <span title="Sat Apr 30 2022 21:20:37 GMT+0200 (heure d’été d’Europe centrale)" aria-label="7 months ago">7mo ago</span></span></span> 
          </span>
          <span class="sc-jIZahH sc-himrzO sc-fXynhf kdTVzc glCpMy ctwKCt">
            <span class="sc-bPPhlf hfaBPJ">
              <a href="/code/pmarcelino/comprehensive-data-exploration-with-python/comments" class="sc-dPyBCJ sc-bBXxYQ sc-bOJcbE cSRCiy cFEurs gTFrUa">1819 comments</a> · 
              <span class="sc-ibQCxQ idHgMS">
                <span class="sc-cKajLJ jNrpDQ">House Prices - Advanced Regression Techniques</span>
              </span>
            </span>
          </span>
        </div>
      </a>
      <div class="sc-gFGZVQ jDMEwY sc-yTtWT kdALiC">
        <div class="sc-dICTr dlQsbO">
          <button mode="default" data-testid="upvotebutton__upvote" aria-label="Upvote" class="sc-dNezTh sc-lkcIho cSGKPD cTyEVx">
            <i class="rmwc-icon rmwc-icon--ligature google-material-icons sc-gKXOVf jWACgA" sizevalue="18px">arrow_drop_up</i>
          </button>
          <span mode="default" class="sc-gXmSlM sc-cCsOjp sc-hAGLhy cKhlzA piYDj mWvOY">12770</span>
        </div>
        <span class="sc-dXxSUK cHoXAr">
          <span class="sc-jIZahH sc-himrzO sc-hRwTwm kdTVzc glCpMy IISDK">
            <img role="presentation" alt="" src="/static/images/medals/competitions/golds@1x.png" style="height: 9px; width: 9px;"> Gold
          </span>
          <div class="mdc-menu-surface--anchor">
            <button aria-label="more_horiz" class="sc-jSMfEi eiMRSN sc-bgrGEg hPFZMI google-material-icons">more_horiz</button>
          </div>
        </span>
      </div>
    </div>
    <div class="sc-lbxAil LkNdN"></div>
  </li>
</ul>

编写的Python代码

import pandas as pd
import requests
from bs4 import BeautifulSoup

headers = {'User-Agent':'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_11_4) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/50.0.2661.94 Safari/537.36'} 


url = "https://www.kaggle.com/code?sortBy=voteCount&page=1"

req = requests.get(url, headers = headers)
soup = BeautifulSoup(req.text, 'html.parser')

html_content = soup.find_all('li', attrs = {'class': 'sc-jfmDQi hfJycS'})  

data = []

for elements in html_content:
    data.append({
   'title': elements.find("div", {"class": "sc-iBkjds sc-fLlhyt sc-fbPSWO uVZhN izULIq A-dENW"}).text,
   'stars': elements.find("span", {"class": "sc-gXmSlM sc-cCsOjp sc-hAGLhy cKhlzA piYDj mWvOY"}).text,
   'resume': elements.find("span", {"class": "sc-cKajLJ jNrpDQ"}).text,
   'comments': elements.find("span", {"class": "sc-dPyBCJ sc-bBXxYQ sc-bOJcbE cSRCiy cFEurs gTFrUa"}).text,
   'link': elements.get('href')})

print(data)

输出结果

[]


问题原因与解决办法

核心问题

  1. 动态内容渲染:Kaggle的代码列表是通过JavaScript动态加载的,requests只能获取页面的静态HTML,无法拿到JS渲染后的实际内容,所以找不到目标元素。
  2. 动态类名不可靠:代码中使用的sc-jfmDQi hfJycS这类类名是前端框架动态生成的,会随时变化,依赖这类类名定位元素会导致失效。
  3. 链接提取错误:目标链接不在<li>标签上,而是在<li>内部的<a>标签中。

解决方案

方案一:使用Selenium获取动态渲染内容

Selenium可以模拟浏览器加载页面,获取完整的渲染后内容,适合处理动态页面:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from bs4 import BeautifulSoup

url = "https://www.kaggle.com/code?sortBy=voteCount&page=1"

# 初始化Chrome浏览器(需提前下载对应版本的ChromeDriver)
driver = webdriver.Chrome()
driver.get(url)

# 等待目标列表加载完成,最多等待10秒
wait = WebDriverWait(driver, 10)
wait.until(EC.presence_of_all_elements_located((By.CSS_SELECTOR, 'ul[role="list"].km-list--three-line li[role="listitem"]')))

# 获取渲染后的页面源码
soup = BeautifulSoup(driver.page_source, 'html.parser')
driver.quit()

data = []
# 基于稳定的属性和结构定位元素,避免依赖动态类名
items = soup.select('ul[role="list"].km-list--three-line li[role="listitem"]')

for item in items:
    # 提取标题
    title_elem = item.select_one('a[href^="/code/"] div:first-of-type')
    title = title_elem.text.strip() if title_elem else None
    
    # 提取点赞数
    stars_elem = item.select_one('button[data-testid="upvotebutton__upvote"] + span')
    stars = stars_elem.text.strip() if stars_elem else None
    
    # 提取关联竞赛
    resume_elem = item.select_one('span[class*="cKajLJ"]')
    resume = resume_elem.text.strip() if resume_elem else None
    
    # 提取评论数
    comments_elem = item.select_one('a[href$="/comments"]')
    comments = comments_elem.text.strip() if comments_elem else None
    
    # 提取完整链接
    link_elem = item.select_one('a[href^="/code/"]')
    link = f"https://www.kaggle.com{link_elem['href']}" if link_elem else None
    
    if title:
        data.append({
            'title': title,
            'stars': stars,
            'resume': resume,
            'comments': comments,
            'link': link
        })

print(data)

方案二:使用Kaggle API(推荐)

Kaggle提供了官方API,可以合规、稳定地获取代码数据,避免反爬问题:

  1. 安装Kaggle库:pip install kaggle
  2. 在Kaggle账号中创建API令牌,保存到本地指定路径
  3. 通过API获取代码列表,示例代码如下:
import kaggle

# 初始化API(需提前配置好Kaggle令牌)
kaggle.api.authenticate()

# 获取按点赞排序的代码列表
codes = kaggle.api.kernels_list(sort_by='voteCount', page=1)

data = []
for code in codes:
    data.append({
        'title': code.title,
        'stars': code.voteCount,
        'resume': code.competitionTitle if hasattr(code, 'competitionTitle') else None,
        'comments': code.commentCount,
        'link': f"https://www.kaggle.com/code/{code.ownerSlug}/{code.slug}"
    })

print(data)

注意事项

  • 使用Selenium时,需确保浏览器驱动版本与浏览器版本匹配,且注意控制请求频率,避免触发Kaggle的反爬机制。
  • 优先选择官方API方案,不仅稳定,还能避免违反Kaggle的使用条款。

内容的提问来源于stack exchange,提问作者ladybug

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 04:55:18