如何使用Python与Tweepy实时提取指定页面的新推文?
Hey there! Let's figure out how to set up real-time tweet extraction for a specific page (I’m assuming you mean either a Twitter user’s timeline or a feed focused on specific keywords/topics) using Python and Tweepy. This solution will catch new tweets the moment they’re posted—no guessing when they’ll go live. Here’s how to do it:
Twitter’s API v2 offers an official Filtered Stream that lets you receive tweets in real time, which is way more efficient than repeatedly checking (polling) for updates. Here’s how to implement it:
Prerequisites First
- You’ll need a Twitter Developer account with an active project, plus your Bearer Token (this is key for API v2 access).
- Install the latest version of Tweepy:
pip install tweepy --upgrade
Step 1: Build a Custom Stream Listener
We’ll create a class that inherits from Tweepy’s StreamingClient to handle incoming tweets. You can customize how you process each new tweet (like printing it, saving to a database, etc.).
Example: Track New Tweets from a Specific User
import tweepy # Replace with your actual Bearer Token BEARER_TOKEN = "YOUR_BEARER_TOKEN_HERE" class RealTimeTweetTracker(tweepy.StreamingClient): def on_response(self, response): # Extract user info and tweet content from the response user = response.includes["users"][0] tweet = response.data print(f"\nNew Tweet from @{user.username}:") print(f"Content: {tweet.text}") print(f"Posted at: {tweet.created_at}") # Add your own logic here—save to a CSV, send a notification, etc. def on_errors(self, errors): # Handle any API errors that pop up print(f"Error encountered: {errors[0]['message']} (Code: {errors[0]['code']})") # Initialize the tracker tweet_tracker = RealTimeTweetTracker(BEARER_TOKEN) # Add a rule to track tweets from your target user (replace with their user ID) # To get a user's ID: Use Tweepy's `client.get_user(username="target_username")` or check Twitter's web interface target_user_id = "123456789" tweet_tracker.add_rules(tweepy.StreamRule(f"from:{target_user_id}")) # Start the stream—we request user info alongside tweets using expansions tweet_tracker.filter(tweet_fields=["created_at"], expansions=["author_id"])
Example: Track Tweets with Specific Keywords
If you want to catch tweets containing certain keywords instead of a user’s timeline, replace the rule with:
tweet_tracker.add_rules(tweepy.StreamRule("your_target_keyword_here"))
Step 2: Manage Stream Rules
You can view, add, or remove rules as needed:
# View all active rules current_rules = tweet_tracker.get_rules() print("Active Rules:", current_rules.data) # Delete existing rules (if needed) if current_rules.data: rule_ids = [rule.id for rule in current_rules.data] tweet_tracker.delete_rules(rule_ids)
If you can’t use the Streaming API (e.g., due to permission limits), you can set up a polling loop that checks for new tweets at regular intervals. This has a slight delay but works reliably.
import tweepy import time BEARER_TOKEN = "YOUR_BEARER_TOKEN_HERE" TARGET_USER_ID = "123456789" CHECK_EVERY_SECONDS = 60 # Adjust this based on how often you want to check # Initialize the Tweepy client client = tweepy.Client(BEARER_TOKEN) # Get the latest tweet ID to start with last_tweet_id = None initial_tweets = client.get_users_tweets(TARGET_USER_ID, max_results=5, tweet_fields=["created_at"]) if initial_tweets.data: last_tweet_id = initial_tweets.data[0].id print("Initial Tweets Loaded:") for tweet in reversed(initial_tweets.data): print(f"- {tweet.text} (Posted: {tweet.created_at})") # Start the polling loop while True: # Fetch only tweets newer than the last one we saw new_tweets = client.get_users_tweets( TARGET_USER_ID, max_results=10, since_id=last_tweet_id, tweet_fields=["created_at"] ) if new_tweets.data: print("\nNew Tweets Found:") # Reverse to process in chronological order for tweet in reversed(new_tweets.data): print(f"- {tweet.text} (Posted: {tweet.created_at})") # Update the last tweet ID to the most recent one last_tweet_id = new_tweets.data[0].id # Wait before checking again time.sleep(CHECK_EVERY_SECONDS)
- Permissions: Make sure your Twitter Developer project has Elevated or Academic Research access to use the Filtered Stream.
- Rate Limits: Always follow Twitter’s API rate limits—polling too frequently can get your access restricted.
- Error Handling: Add extra logic to handle network drops (e.g., auto-reconnect the stream) or API errors.
- Persistence: If you need to store tweets long-term, write them to a CSV/JSON file or a database like SQLite or PostgreSQL.
内容的提问来源于stack exchange,提问作者Bilal Aamir

