使用R的rvest爬取Kaggle挑战链接时遇空节点问题
Hey Nickkon, let's work through this Kaggle scraping issue together. I've run into similar headaches with brittle XPath selectors and dynamic content before—here's how to get back on track:
1. Why Your Current XPath Isn't Working
Absolute XPaths like /html/body/div[1]/div[2]/... are super fragile. Kaggle regularly updates its page structure, and many of its content sections load dynamically (via JavaScript) after the initial page load. That's why your deep-div targeting is returning empty content—either the structure shifted, or the content hasn't loaded when rvest tries to scrape it.
2. Switch to Stable, Relative Selectors
Instead of relying on strict page hierarchy, use selectors that target consistent attributes or class names. Kaggle's competition links all follow a predictable pattern: their href starts with /competitions/. This makes attribute-based selectors way more reliable.
Here's a revised rvest script to grab the first page of competition links and titles:
library(rvest) library(dplyr) # Target the main competitions page kaggle_comp_url <- "https://www.kaggle.com/competitions" page <- read_html(kaggle_comp_url) # Grab all competition links (using href pattern for stability) competition_links <- page %>% html_elements("[href^='/competitions/']") %>% html_attr("href") %>% paste0("https://www.kaggle.com", .) %>% # Convert relative URLs to full ones unique() # Remove duplicate links (some cards have multiple links) # Grab corresponding titles competition_titles <- page %>% html_elements("[href^='/competitions/']") %>% html_text2() %>% unique() # Combine into a tidy data frame competition_df <- tibble( title = competition_titles, link = competition_links ) # Check the results head(competition_df)
3. Handling Dynamic Loaded Content (For All Competitions)
If you want more than just the first page, rvest alone won't cut it—Kaggle loads additional competitions as you scroll. You'll need to simulate a browser to trigger the JavaScript loading. Here's a quick example using RSelenium:
library(RSelenium) library(rvest) library(dplyr) # Start a Chrome browser instance (make sure ChromeDriver is installed) driver <- rsDriver(browser = "chrome", port = 4567L) remDr <- driver[["client"]] # Navigate to Kaggle competitions remDr$navigate("https://www.kaggle.com/competitions") # Simulate scrolling to load more content (adjust loop count based on how many you need) for (i in 1:5) { remDr$executeScript("window.scrollTo(0, document.body.scrollHeight);") Sys.sleep(3) # Give time for content to load } # Extract the fully loaded page source page_source <- remDr$getPageSource()[[1]] page <- read_html(page_source) # Use the same selectors as before to grab data competition_links <- page %>% html_elements("[href^='/competitions/']") %>% html_attr("href") %>% paste0("https://www.kaggle.com", .) %>% unique() competition_titles <- page %>% html_elements("[href^='/competitions/']") %>% html_text2() %>% unique() competition_df <- tibble(title = competition_titles, link = competition_links) # Clean up: close the browser remDr$close() driver$server$stop()
4. Pro Tip for Testing Selectors
Open Kaggle's competitions page, hit F12 to open DevTools, and use the element picker to click a competition card. Look for consistent class names or attributes (like the href pattern) that you can target—this helps you build selectors that won't break when Kaggle updates its UI.
内容的提问来源于stack exchange,提问作者Nickkon

