如何通过Tweepy按小时获取指定时间范围的Twitter推文?
Hey there! Let's figure out how to fetch hourly tweets using Tweepy, since the Twitter API's since and until parameters only accept date strings (not full datetime with hours). Here are two reliable approaches to make this work:
方法1:本地过滤推文时间
First, we'll fetch all tweets from the target date, then filter them locally to keep only those within your desired hourly window. This is great if you need flexibility with time ranges or want to process multiple hours at once.
Code Example
import tweepy import datetime import pytz # Assume you've already set up your API authentication here auth = tweepy.OAuthHandler(consumer_key, consumer_secret) auth.set_access_token(access_token, access_token_secret) api = tweepy.API(auth) search_words = 'bitcoin -filter:retweet' # Grab all tweets from the full target date date_since = '2019-08-19' date_until = '2019-08-20' # Define your target hourly window (use UTC, since Twitter uses UTC for tweet timestamps) target_start = datetime.datetime(2019, 8, 19, 17, 0, 0, tzinfo=pytz.utc) target_end = datetime.datetime(2019, 8, 19, 18, 0, 0, tzinfo=pytz.utc) hourly_tweets = [] # Iterate through tweets and filter by time for tweet in tweepy.Cursor(api.search, q=search_words, lang="en", since=date_since, until=date_until).items(): tweet_time = tweet.created_at # This is a timezone-aware UTC datetime object if target_start <= tweet_time < target_end: hourly_tweets.append(tweet) # Verify the result print(f"Found {len(hourly_tweets)} tweets in the 2019-08-19 17:00-18:00 UTC window")
Note: If your target time is in a local timezone (not UTC), convert it to UTC first using pytz to avoid mismatches with Twitter's timestamps.
方法2:使用Twitter高级搜索语法直接过滤
Twitter's Search API supports precise time ranges directly in your search query, which is more efficient because the API returns only the tweets you need (no local filtering required).
Code Example
import tweepy # Assume you've already set up your API authentication here auth = tweepy.OAuthHandler(consumer_key, consumer_secret) auth.set_access_token(access_token, access_token_secret) api = tweepy.API(auth) # Build your search query with the exact hourly time range # Use the format `since:YYYY-MM-DD_HH:mm:ss` and `until:YYYY-MM-DD_HH:mm:ss` search_words = 'bitcoin -filter:retweet since:2019-08-19_17:00:00 until:2019-08-19_18:00:00' # Fetch tweets directly from the API with the time range included in the query tweets = tweepy.Cursor(api.search, q=search_words, lang="en").items(100) # Process your hourly tweets for tweet in tweets: print(f"[{tweet.created_at}] {tweet.text[:50]}...")
Note: This method relies on Twitter's search indexing, so there might be minor delays, but it's perfect for most use cases where you want to minimize local processing.
Quick Comparison
- Method 1 is flexible for complex time filtering but requires fetching more tweets upfront.
- Method 2 is more efficient and reduces data transfer, but is tied to Twitter's search syntax limitations.
内容的提问来源于stack exchange,提问作者Mert Yanık

