使用Tweepy调用Twitter API时写入文本文件失败求助
Hey there! Let's work through this Tweepy data fetching and writing issue step by step. Since you've already got your authentication sorted and have tweet IDs loaded into a DataFrame, the problem usually comes down to mishandling API responses, incorrect data processing, or small oversights in file writing. Here's how to fix it:
Make sure you're using the correct endpoint for your Twitter API version (v2 is recommended now). Let's start with a solid, tested fetching snippet:
import tweepy import pandas as pd # Your existing auth setup (fill in your tokens) client = tweepy.Client( bearer_token=YOUR_BEARER_TOKEN, consumer_key=YOUR_API_KEY, consumer_secret=YOUR_API_SECRET, access_token=YOUR_ACCESS_TOKEN, access_token_secret=YOUR_ACCESS_TOKEN_SECRET, wait_on_rate_limit=True # Auto-handle rate limits ) # Clean and prepare tweet IDs from your DataFrame # (Remove NaNs, convert to strings to avoid type issues) tweet_ids = df['tweet_id'].dropna().astype(str).tolist() # Batch fetch tweets (API allows max 100 IDs per call) response = client.get_tweets( ids=tweet_ids, tweet_fields=['created_at', 'public_metrics', 'author_id'] # Add fields you need ) # Check if data was returned if response.data: print(f"Fetched {len(response.data)} valid tweets") else: print("No tweets returned—check if IDs are valid (not deleted/banned)") # Optional: Check for errors with invalid IDs if response.errors: print("Invalid tweet IDs found:") for error in response.errors: print(f"ID {error['value']}: {error['detail']}")
Convert the raw Tweepy response into a format that's easy to turn into a DataFrame and write to a file:
# Build a list of dictionaries with your desired data points processed_tweets = [] for tweet in response.data: tweet_info = { 'tweet_id': tweet.id, 'text': tweet.text, 'created_at': tweet.created_at, 'retweet_count': tweet.public_metrics['retweet_count'], 'like_count': tweet.public_metrics['like_count'], 'author_id': tweet.author_id } processed_tweets.append(tweet_info) # Create your new DataFrame new_df = pd.DataFrame(processed_tweets)
Avoid common pitfalls like encoding issues or permission errors with these snippets:
Write to Text File (JSON Lines Format)
import json # Use utf-8 encoding to handle emojis/special characters with open('tweets_output.txt', 'w', encoding='utf-8') as f: for tweet in processed_tweets: json.dump(tweet, f, ensure_ascii=False) f.write('\n') # Each tweet on a new line for readability
Save the New DataFrame
# Save to CSV (or Excel if preferred) new_df.to_csv('updated_tweets_df.csv', index=False, encoding='utf-8') # Or if you want a pickled DataFrame for later Python use new_df.to_pickle('updated_tweets_df.pkl')
- Missing Data: If fewer tweets are returned than your input IDs, check
response.errorsto see which IDs are invalid (deleted tweets, suspended accounts, etc.). - Rate Limits: The
wait_on_rate_limit=Trueparameter in the Client setup will make Tweepy automatically pause when you hit API limits—no need to handle retries manually. - File Permissions: Ensure you're writing to a directory you have access to (avoid system-protected folders like
C:\Windowsor/root). - Data Type Issues: Double-check that your
tweet_idcolumn in the original DataFrame has no blank values or non-numeric/non-string entries—usedf['tweet_id'].dropna().astype(str)to clean it up.
内容的提问来源于stack exchange,提问作者Jw007

