You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用R的rvest爬取Kaggle挑战链接时遇空节点问题

Hey Nickkon, let's work through this Kaggle scraping issue together. I've run into similar headaches with brittle XPath selectors and dynamic content before—here's how to get back on track:

1. Why Your Current XPath Isn't Working

Absolute XPaths like /html/body/div[1]/div[2]/... are super fragile. Kaggle regularly updates its page structure, and many of its content sections load dynamically (via JavaScript) after the initial page load. That's why your deep-div targeting is returning empty content—either the structure shifted, or the content hasn't loaded when rvest tries to scrape it.

2. Switch to Stable, Relative Selectors

Instead of relying on strict page hierarchy, use selectors that target consistent attributes or class names. Kaggle's competition links all follow a predictable pattern: their href starts with /competitions/. This makes attribute-based selectors way more reliable.

Here's a revised rvest script to grab the first page of competition links and titles:

library(rvest)
library(dplyr)

# Target the main competitions page
kaggle_comp_url <- "https://www.kaggle.com/competitions"
page <- read_html(kaggle_comp_url)

# Grab all competition links (using href pattern for stability)
competition_links <- page %>%
  html_elements("[href^='/competitions/']") %>%
  html_attr("href") %>%
  paste0("https://www.kaggle.com", .) %>% # Convert relative URLs to full ones
  unique() # Remove duplicate links (some cards have multiple links)

# Grab corresponding titles
competition_titles <- page %>%
  html_elements("[href^='/competitions/']") %>%
  html_text2() %>%
  unique()

# Combine into a tidy data frame
competition_df <- tibble(
  title = competition_titles,
  link = competition_links
)

# Check the results
head(competition_df)

3. Handling Dynamic Loaded Content (For All Competitions)

If you want more than just the first page, rvest alone won't cut it—Kaggle loads additional competitions as you scroll. You'll need to simulate a browser to trigger the JavaScript loading. Here's a quick example using RSelenium:

library(RSelenium)
library(rvest)
library(dplyr)

# Start a Chrome browser instance (make sure ChromeDriver is installed)
driver <- rsDriver(browser = "chrome", port = 4567L)
remDr <- driver[["client"]]

# Navigate to Kaggle competitions
remDr$navigate("https://www.kaggle.com/competitions")

# Simulate scrolling to load more content (adjust loop count based on how many you need)
for (i in 1:5) {
  remDr$executeScript("window.scrollTo(0, document.body.scrollHeight);")
  Sys.sleep(3) # Give time for content to load
}

# Extract the fully loaded page source
page_source <- remDr$getPageSource()[[1]]
page <- read_html(page_source)

# Use the same selectors as before to grab data
competition_links <- page %>%
  html_elements("[href^='/competitions/']") %>%
  html_attr("href") %>%
  paste0("https://www.kaggle.com", .) %>%
  unique()

competition_titles <- page %>%
  html_elements("[href^='/competitions/']") %>%
  html_text2() %>%
  unique()

competition_df <- tibble(title = competition_titles, link = competition_links)

# Clean up: close the browser
remDr$close()
driver$server$stop()

4. Pro Tip for Testing Selectors

Open Kaggle's competitions page, hit F12 to open DevTools, and use the element picker to click a competition card. Look for consistent class names or attributes (like the href pattern) that you can target—this helps you build selectors that won't break when Kaggle updates its UI.

内容的提问来源于stack exchange,提问作者Nickkon

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 10:20:01