You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python实现推特数据JSON转CSV:解决属性匹配与代码适配问题

Fixing Twitter Data Export to CSV with twitterscraper

Let's break down what's going wrong with your code and fix it step by step:

Key Issues in Your Current Code

  • Incorrect traversal of all_tweets: Your all_tweets is a dictionary where keys are usernames and values are lists of tweet objects. When you loop for tweet in all_tweets, you're iterating over usernames (strings), not the actual tweet objects—this is why you're getting the "invalid attribute" error.
  • Mismatched attribute names: You used fullname_str but the actual tweet object attribute (and your JSON structure) uses fullname. Also, encoding tweet.text to bytes will create messy output in your CSV; keep it as a string instead.

Corrected Working Code

Here's the revised version that properly extracts all tweets and saves them to a clean CSV:

from twitterscraper import query_tweets
import csv
import datetime as dt

# Load username list
with open('account_list.txt', 'r') as f:
    handles = f.readlines()
    handles = [h.rstrip('\n') for h in handles]

all_tweets = {}

if __name__ == '__main__':
    # Fetch tweets for each user
    for handle in handles:
        user_tweets = query_tweets(
            f"from:{handle}",
            limit=10000,
            begindate=dt.date(2010, 1, 1),
            enddate=dt.date.today(),
            poolsize=10,
            lang=''
        )
        all_tweets[handle] = user_tweets
        print(f"Completed fetching tweets for {handle}")

    # Flatten all tweets into a single list (one row per tweet)
    outtweets = []
    for username, tweet_list in all_tweets.items():
        for tweet in tweet_list:
            outtweets.append([
                tweet.fullname,
                tweet.id,
                tweet.likes,
                tweet.replies,
                tweet.retweets,
                tweet.text,
                tweet.timestamp,
                tweet.url,
                username  # Use tweet.user if that attribute is available and preferred
            ])

    # Write to CSV with proper encoding
    with open('all_tweets.csv', 'w', newline='', encoding='utf-8') as f:
        writer = csv.writer(f)
        # Write header row
        writer.writerow([
            "fullname", "id", "likes", "replies", "retweets",
            "text", "timestamp", "url", "user"
        ])
        # Write all tweet data
        writer.writerows(outtweets)

What We Changed

  1. Properly traversed tweet data: We loop through each user in all_tweets, then each tweet in their list—this ensures we're accessing actual tweet objects, not just usernames.
  2. Fixed attribute names: Swapped fullname_str for fullname and removed the unnecessary .encode('utf-8') on tweet.text to preserve readable text.
  3. Improved CSV handling: Added encoding='utf-8' to support special characters in tweets, and newline='' to avoid extra blank rows on Windows systems.
  4. Made code more debuggable: Used explicit loops instead of a single dense list comprehension, which will make it easier to troubleshoot issues with your large dataset of 8000 users.

How to Verify Tweet Object Attributes

If you're ever unsure about the exact attributes available on a tweet object, add this quick check after fetching tweets for a test user:

# Add this inside the handle loop, after fetching user_tweets
if user_tweets:
    print(dir(user_tweets[0]))

This will print all available attributes for the tweet object—match these to your CSV columns to avoid attribute errors.

内容的提问来源于stack exchange,提问作者GreenPirate

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 08:44:37