如何使用R基于国会议员列表高效批量提取Twitter推文?
Absolutely! You can absolutely use your prepped list of congressional members to batch-extract tweets with rtweet—no need to manually input each user ID or screen name. Here's a step-by-step, efficient approach tailored for 600+ users, including handling API rate limits and potential errors:
First, make sure you have these packages installed and loaded. dplyr helps with data manipulation, purrr simplifies batch operations, and rtweet is your core tool for Twitter data:
# Install packages if you haven't already install.packages(c("rtweet", "dplyr", "purrr")) # Load packages library(rtweet) library(dplyr) library(purrr)
Assuming you’ve saved your congressional members list as a CSV (with a column for either screen_name or user_id), import it into R. Let’s say your file is named congress_members.csv:
# Import the CSV file congress_users <- read.csv("congress_members.csv", stringsAsFactors = FALSE) # Extract the vector of user identifiers (use user_id if you have that instead) user_list <- congress_users$screen_name
Twitter’s API has strict rate limits (for get_timeline, you can make 900 requests every 15 minutes). To avoid hitting limits and handle cases where accounts are suspended/deleted, we’ll:
- Split users into small batches
- Use a "safe" version of
get_timelinethat skips problematic accounts - Add delays between batches to respect rate limits
Here’s the code:
# Create a "safe" version of get_timeline that returns NULL instead of failing on errors safe_get_timeline <- possibly(get_timeline, otherwise = NULL) # Split your user list into batches (adjust batch_size based on rate limits) batch_size <- 50 # 50 users per batch keeps you under the 900-request/15min limit user_batches <- split(user_list, ceiling(seq_along(user_list) / batch_size)) # Loop through each batch to extract tweets all_congress_tweets <- map_dfr(user_batches, function(batch) { # Extract tweets for the current batch (n = 3200 is the max per user) batch_tweets <- map_dfr(batch, ~safe_get_timeline(.x, n = 3200)) # Pause for 60 seconds between batches to avoid triggering rate limits Sys.sleep(60) return(batch_tweets) })
Once extraction is complete, save the combined tweets to a CSV for later analysis:
write.csv(all_congress_tweets, "congressional_tweets.csv", row.names = FALSE)
- API Authentication: Before running any code, make sure you’ve set up your Twitter API credentials with
auth_setup_default()(follow rtweet’s prompts to link your developer account). - Adjust Batch Size/Delay: If you hit rate limits, try reducing the batch size or increasing the
Sys.sleep()duration (e.g., 90 seconds instead of 60). - User IDs vs Screen Names: If your list uses
user_idinstead ofscreen_name, just replacecongress_users$screen_namewithcongress_users$user_id—get_timelineworks with both. - Check for Missing Data: After extraction, you can use
filter(all_congress_tweets, is.na(screen_name))to spot any users that couldn’t be processed.
内容的提问来源于stack exchange,提问作者Golnar Teimouri

