Python实现推特数据JSON转CSV:解决属性匹配与代码适配问题
Fixing Twitter Data Export to CSV with twitterscraper
Let's break down what's going wrong with your code and fix it step by step:
Key Issues in Your Current Code
- Incorrect traversal of
all_tweets: Yourall_tweetsis a dictionary where keys are usernames and values are lists of tweet objects. When you loopfor tweet in all_tweets, you're iterating over usernames (strings), not the actual tweet objects—this is why you're getting the "invalid attribute" error. - Mismatched attribute names: You used
fullname_strbut the actual tweet object attribute (and your JSON structure) usesfullname. Also, encodingtweet.textto bytes will create messy output in your CSV; keep it as a string instead.
Corrected Working Code
Here's the revised version that properly extracts all tweets and saves them to a clean CSV:
from twitterscraper import query_tweets import csv import datetime as dt # Load username list with open('account_list.txt', 'r') as f: handles = f.readlines() handles = [h.rstrip('\n') for h in handles] all_tweets = {} if __name__ == '__main__': # Fetch tweets for each user for handle in handles: user_tweets = query_tweets( f"from:{handle}", limit=10000, begindate=dt.date(2010, 1, 1), enddate=dt.date.today(), poolsize=10, lang='' ) all_tweets[handle] = user_tweets print(f"Completed fetching tweets for {handle}") # Flatten all tweets into a single list (one row per tweet) outtweets = [] for username, tweet_list in all_tweets.items(): for tweet in tweet_list: outtweets.append([ tweet.fullname, tweet.id, tweet.likes, tweet.replies, tweet.retweets, tweet.text, tweet.timestamp, tweet.url, username # Use tweet.user if that attribute is available and preferred ]) # Write to CSV with proper encoding with open('all_tweets.csv', 'w', newline='', encoding='utf-8') as f: writer = csv.writer(f) # Write header row writer.writerow([ "fullname", "id", "likes", "replies", "retweets", "text", "timestamp", "url", "user" ]) # Write all tweet data writer.writerows(outtweets)
What We Changed
- Properly traversed tweet data: We loop through each user in
all_tweets, then each tweet in their list—this ensures we're accessing actual tweet objects, not just usernames. - Fixed attribute names: Swapped
fullname_strforfullnameand removed the unnecessary.encode('utf-8')ontweet.textto preserve readable text. - Improved CSV handling: Added
encoding='utf-8'to support special characters in tweets, andnewline=''to avoid extra blank rows on Windows systems. - Made code more debuggable: Used explicit loops instead of a single dense list comprehension, which will make it easier to troubleshoot issues with your large dataset of 8000 users.
How to Verify Tweet Object Attributes
If you're ever unsure about the exact attributes available on a tweet object, add this quick check after fetching tweets for a test user:
# Add this inside the handle loop, after fetching user_tweets if user_tweets: print(dir(user_tweets[0]))
This will print all available attributes for the tweet object—match these to your CSV columns to avoid attribute errors.
内容的提问来源于stack exchange,提问作者GreenPirate
相关产品推荐
相关产品推荐

