rtweet包推文下载:兼顾R Studio任务、断点续传及抽样方案问询
Hey there! Let's tackle your two rtweet questions one by one—they're both super practical for long-term tweet collection projects like yours.
First off, you’re right that the default stream_tweets() blocks your R session, but there are two key fixes here: running the stream in the background so you can keep using RStudio, and setting up your workflow to resume from where you left off if the stream gets interrupted.
Run the Stream in the Background (No More Blocked Sessions!)
Instead of running the stream directly in your main RStudio window, use the callr package to spin up a separate background process. This lets you keep working on other tasks while the stream runs quietly in the background. Here’s how:
library(callr) # Launch the stream in a background process bg_stream <- r_bg( func = function() { library(rtweet) # Stream tweets with "corruption" and save to a JSON file stream_tweets( q = "corruption", file_name = "corruption_tweets.json", append = TRUE, # Append new tweets instead of overwriting timeout = 604800 # Run for 1 week (604800 seconds) ) } ) # Check if the stream is still running bg_stream$status() # Stop the stream manually if needed bg_stream$kill()
You can also wrap this in a loop if you want it to restart automatically after each week, but even if it stops for any reason, we can resume from the last collected tweet.
Resume From a Breakpoint
To avoid re-downloading tweets you already have, use the since_id parameter to tell rtweet to only fetch tweets posted after the most recent one in your existing dataset:
library(rtweet) # Load the tweets you've already collected existing_tweets <- parse_stream("corruption_tweets.json") # Grab the ID of the newest tweet in your dataset latest_tweet_id <- max(existing_tweets$status_id) # Restart the stream, starting after that latest ID stream_tweets( q = "corruption", file_name = "corruption_tweets.json", append = TRUE, since_id = latest_tweet_id, timeout = 604800 )
Pro tip: If your stream is down for more than a few days, you might miss some tweets that fall outside Twitter’s search window. For gaps longer than a week, use search_tweets() with since_id and until_id to fill in the missing data before restarting the stream.
Setting up automated random sampling is totally doable—you’ll just need to schedule a separate script to run at regular intervals (e.g., weekly) to pull a random subset of your collected tweets.
Step 1: Write a Sampling Script
Create a standalone R script (let’s call it sample_corruption_tweets.R) with this code:
library(rtweet) library(dplyr) # Load all collected tweets all_tweets <- parse_stream("corruption_tweets.json") # Randomly sample 10% of the tweets (adjust the fraction as needed) sampled_tweets <- sample_frac(all_tweets, 0.1) # Save the sample with a timestamp to avoid overwriting saveRDS( sampled_tweets, paste0("corruption_tweets_sample_", Sys.Date(), ".rds") )
Step 2: Schedule the Script to Run Automatically
Use a task scheduler to run this script on your preferred schedule:
- Windows: Use the
taskscheduleRpackage to set up weekly (or daily) tasks:library(taskscheduleR) taskscheduler_create( taskname = "SampleCorruptionTweets", rscript = "path/to/your/sample_corruption_tweets.R", schedule = "weekly", starttime = "09:00", days = "Monday" # Pick your preferred day ) - Mac/Linux: Use the
cronRpackage to set up a cron job, which works similarly for scheduling recurring tasks.
Bonus: In-Stream Sampling (If You Want Real-Time Samples)
If you want to sample tweets while the stream is running (instead of waiting to sample the full dataset), you can use the later package to trigger sampling at set intervals without interrupting the stream. Here’s a quick example:
library(rtweet) library(later) # Define a sampling function that runs every 24 hours sample_on_interval <- function() { all_tweets <- parse_stream("corruption_tweets.json") sampled <- sample_frac(all_tweets, 0.1) saveRDS(sampled, paste0("daily_sample_", Sys.Date(), ".rds")) # Schedule the next run in 24 hours later(sample_on_interval, 86400) } # Start the stream and the sampling scheduler stream_tweets( q = "corruption", file_name = "corruption_tweets.json", append = TRUE, timeout = Inf # Run indefinitely (use with caution!) ) sample_on_interval()
Just note that running the stream with timeout = Inf means it’ll run until you manually stop it, so make sure your computer stays on (or use a server).
内容的提问来源于stack exchange,提问作者Vrooshi

