You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python Tweepy调用Twitter Streaming API时如何规避420错误?

Hey there, let's work through your Tweepy Streaming API issues step by step—both the missing data and that annoying 420 rate limit error.

First: Fix the Immediate "No Data" Problem

Your code has a huge red flag right now: you're calling myStream.disconnect() right after starting the stream with async=True. Since async=True runs the stream in a background thread, your main thread immediately jumps to disconnecting the stream before it even has a chance to fetch any tweets. That's why you're seeing zero output!

Fix options:

  • Drop the async=True: Let the stream run in the main thread—this will keep your program running until you manually stop it (like pressing Ctrl+C).
  • Add a wait if you need async: If you must run the stream in the background, add a loop or delay to keep the main thread alive (e.g., time.sleep(3600) to run for an hour).

Next: Beat the 420 Rate Limit Error

A 420 error means Twitter is throttling your requests because you're hitting rate limits or triggering anti-abuse mechanisms (like frequent reconnects). Here's how to fix it:

  • Add exponential backoff to your error handler: When you hit 420, wait longer each time before retrying—this tells Twitter you're not spamming requests.
  • Avoid unnecessary reconnects: Fixing the disconnect() issue above will stop you from repeatedly starting/stopping the stream, which is a common trigger for 420s.
  • Narrow your filter: If you're tracking too many keywords or broad terms, you might be sending too much traffic. Try reducing the number of tracked terms or using more specific filters.
  • Check your API permissions: Make sure your developer account has access to the v1.1 Streaming API (or switch to v2's Filtered Stream if you're on the newer API, since v1.1 is being phased out).

Revised Working Code

Here's your code with all the fixes applied, including exponential backoff for 420 errors and proper stream management:

import json
import time
import tweepy

# Replace with your actual credentials
access_token = 'XXXXXX'
access_token_secret = 'XXXXXX'
consumer_key = 'XXXXXX'
consumer_secret = 'XXXXXX'

class MyListener(tweepy.StreamListener):
    def __init__(self):
        super().__init__()
        self.retry_count = 0  # Track retries for backoff

    def on_status(self, status):
        # Print clean tweet text (handles retweets/extended text automatically)
        print(status.text)
        return True

    def on_data(self, tweet_data):
        # Use this if you need raw JSON instead of parsed status objects
        try:
            data = json.loads(tweet_data)
            print(json.dumps(data, indent=2))  # Pretty-print JSON for readability
        except json.JSONDecodeError:
            print("Failed to parse tweet data")
        return True

    def on_error(self, status_code):
        print(f"Error {status_code} encountered")
        # Handle 420 rate limit with exponential backoff
        if status_code == 420:
            wait_time = 10 * (2 ** self.retry_count)
            print(f"Rate limited! Waiting {wait_time} seconds before reconnecting...")
            time.sleep(wait_time)
            self.retry_count += 1
            return True  # Tell Tweepy to keep trying to reconnect
        # For other errors, return False to disconnect the stream
        return False

# Initialize authentication
auth = tweepy.OAuthHandler(consumer_key=consumer_key, consumer_secret=consumer_secret)
auth.set_access_token(access_token, access_token_secret)

# Set up stream
listener = MyListener()
stream = tweepy.Stream(auth=auth, listener=listener)

# Start streaming (runs in main thread until Ctrl+C)
try:
    print("Starting stream...")
    stream.filter(languages=['en'], track=['@NBA'], async=False)
except KeyboardInterrupt:
    print("\nStopping stream...")
    stream.disconnect()
    print("Stream disconnected successfully")

Quick Extra Tips

  • Choose either on_status or on_data: You don't need both—on_status parses tweets into convenient objects, while on_data gives you raw JSON. Pick whichever fits your use case.
  • Secure your credentials: Never hardcode API keys in your code. Use environment variables or a secure config file instead.
  • Consider Twitter API v2: The v1.1 Streaming API is being deprecated. For newer projects, use Tweepy's StreamingClient to work with v2's Filtered Stream—it's more powerful and supported long-term.

内容的提问来源于stack exchange,提问作者tushaR

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 06:33:06