如何在R中为多个Twitter数据框添加用户粉丝数并封装成函数?
Hey there! Great job getting the workflow working for a single artist's tweets—let's wrap that logic into a reusable function so you can process all 6 of your data frames in one go. Here's how to do it step by step:
Step 1: Create a Reusable Function
First, let's turn your existing code into a function that takes a tweet data frame and returns the same data frame with a new followersCount column. This function will handle all the lookup, matching, and column addition automatically:
library(twitteR) # Define the function to add followers count to a tweet data frame add_followers_count <- function(tweets_df) { # Extract unique screen names to avoid redundant API calls unique_screens <- unique(tweets_df$screenName) # Look up user information for these screen names users <- lookupUsers(unique_screens) # Convert user data to a data frame users_df <- twListToDF(users) # Match each tweet's screen name to the corresponding follower count # We use match() to align the order correctly tweets_df$followersCount <- users_df$followersCount[match(tweets_df$screenName, users_df$screenName)] # Return the updated data frame return(tweets_df) }
Quick notes on the function:
- We first grab
unique_screensto cut down on unnecessary API calls (since multiple tweets might come from the same user)—this makes the process faster and helps avoid hitting Twitter's rate limits. - The
match()function ensures each tweet gets the correct follower count by aligningscreenNamevalues between your tweet data and the user data we pull.
Step 2: Batch Process All Your Data Frames
Now, let's group all 6 of your tweet data frames into a list (this is the simplest way to handle batch tasks in R). Let's assume your data frames are named tweets_artist1_df, tweets_artist2_df, ..., tweets_artist6_df:
# Create a list containing all your tweet data frames tweets_list <- list( artist1 = tweets_artist1_df, artist2 = tweets_artist2_df, artist3 = tweets_artist3_df, artist4 = tweets_artist4_df, artist5 = tweets_artist5_df, artist6 = tweets_artist6_df ) # Apply our function to every data frame in the list updated_tweets_list <- lapply(tweets_list, add_followers_count) # If you want to pull individual updated data frames back out of the list: tweets_artist1_df_updated <- updated_tweets_list$artist1 tweets_artist2_df_updated <- updated_tweets_list$artist2 # ... repeat for the remaining artists
Bonus: Add a Safety Net for API Rate Limits
Twitter's API has rate limits, so if you're dealing with a huge number of unique users, you might hit a block. To avoid this, we can add a short delay before each API call. Here's an updated version of the function:
add_followers_count_safe <- function(tweets_df, delay = 5) { unique_screens <- unique(tweets_df$screenName) # Add a delay before the API call to respect rate limits Sys.sleep(delay) users <- lookupUsers(unique_screens) users_df <- twListToDF(users) tweets_df$followersCount <- users_df$followersCount[match(tweets_df$screenName, users_df$screenName)] return(tweets_df) } # Use this safe version for batch processing updated_tweets_list_safe <- lapply(tweets_list, add_followers_count_safe)
This should handle all your batch processing needs smoothly—you'll end up with every data frame having the followersCount column added correctly.
内容的提问来源于stack exchange,提问作者tivoo

