使用rvest抓取Twitter账号推文数量失效,寻求解决建议
我之前用这段R代码抓取Twitter账号的推文数量(比API方式更快):
library("rvest") scrape = read_html("https://twitter.com/whitehouse/") scrape.count = scrape %>% html_nodes(".js-nav .ProfileNav-value") %>% html_text() %>% as.character() scrape.count = as.numeric(paste(substr(scrape.count,1,1),substr(scrape.count,3,5),sep=""))这段代码之前完全正常,但两天前开始返回空值。我用SelectorGadget确认过选择器没变化,也知道rvest在动态网站上可能有问题,但之前一直没异常。有没有解决建议?
Hey there, let's break down what's happening and how to get your scraper working again:
Why This Stopped Working
Twitter (now rebranded as X) regularly adjusts how it serves page content, even if CSS selectors appear unchanged. The most probable cause is that the tweet count is now dynamically loaded via JavaScript—the initial static HTML that read_html() fetches no longer includes this data. It's also possible Twitter added basic anti-scraping checks that block non-browser requests.
Fix 1: Use Browser Automation for Dynamic Content
Since rvest can't execute JavaScript, you'll need to simulate a real browser to load the full, rendered page. RSelenium or the playwright R package are solid options. Here's an RSelenium example:
library(RSelenium) library(rvest) # Launch Chrome (ensure ChromeDriver is installed and in your system PATH) driver <- rsDriver(browser = "chrome", port = 4567L, verbose = FALSE) remDr <- driver$client # Navigate to the page and wait for dynamic content to load remDr$navigate("https://twitter.com/whitehouse/") Sys.sleep(3) # Adjust delay based on your internet speed # Grab the fully rendered page source page_source <- remDr$getPageSource()[[1]] scrape <- read_html(page_source) # Extract and clean the tweet count scrape.count <- scrape %>% html_nodes(".js-nav .ProfileNav-value") %>% html_text() %>% as.character() %>% gsub("[^0-9]", "", .) %>% # Remove non-numeric characters like commas as.numeric() # Clean up browser session remDr$close() driver$server$stop()
Fix 2: Mimic Browser Requests with Custom Headers
Twitter might be blocking requests that don't look like they're coming from a real browser. Try adding a user-agent header using httr:
library(rvest) library(httr) # Mimic a Chrome browser request headers <- add_headers( "User-Agent" = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" ) # Fetch the page with headers response <- GET("https://twitter.com/whitehouse/", headers) scrape <- read_html(response) # Attempt to extract the count again scrape.count <- scrape %>% html_nodes(".js-nav .ProfileNav-value") %>% html_text() %>% as.character()
Note: This fix might stop working quickly as Twitter updates its anti-scraping rules.
Fix 3: Extract Data from Embedded Page JSON
Many dynamic sites store core content in inline JSON scripts. Twitter includes user data in an INITIAL_STATE script tag—you can parse this directly:
library(rvest) library(jsonlite) scrape <- read_html("https://twitter.com/whitehouse/") # Extract the initial state JSON blob script_data <- scrape %>% html_nodes("script[type='application/json'][data-testid='INITIAL_STATE']") %>% html_text() # Parse JSON and pull the tweet count initial_state <- fromJSON(script_data) tweet_count <- initial_state$entities$users$`whitehouse`$statuses_count
This is more reliable than CSS selectors because it pulls directly from the data the site uses to render the page—just double-check the JSON path if Twitter updates its structure later.
Fix 4: Switch to the Official API (Long-Term Solution)
While you mentioned API calls are slower, the Twitter API v2 (via the rtweet package) is far more stable than scraping. It won't break when Twitter tweaks its page layout:
library(rtweet) # Set up authentication (you'll need to create an app on X's Developer Platform) auth_setup_default() # Fetch user data and extract the tweet count user_info <- lookup_users("whitehouse") tweet_count <- user_info$statuses_count
Final Notes
The core issue is almost certainly that the tweet count no longer exists in the static HTML response. For quick fixes, go with browser automation or JSON extraction. For a sustainable solution, invest in setting up the official API—it'll save you from constant troubleshooting as Twitter updates its site.
内容的提问来源于stack exchange,提问作者David Antoš

