You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

新手求助:无法处理Archive.org的Twitter JSON数据集

Hey there! No worries at all—everyone starts somewhere, and handling massive Twitter JSON datasets can be tricky even for seasoned developers. Let's walk through how to fix your issues step by step.

First, let's clarify what you're working with: those Archive.org Twitter files are JSON Lines (.jsonl) files, meaning each line is a separate JSON object (one Tweet per line). That's why trying to open the whole file at once crashes your tools—it's way too big to fit in memory. We'll use line-by-line processing to get around that.

Step 1: Read the File & Filter by Hashtag

Instead of loading the entire file into memory, we'll read one line at a time, parse the Tweet, and check if it contains your target hashtag. Here's a simple Python script to do that:

import json
from datetime import datetime

# Customize these variables to match your needs
TARGET_HASHTAG = "#YourHashtagHere"  # Replace with your desired hashtag
INPUT_FILE_PATH = "path/to/your/twitter-file.json"  # Update this path
OUTPUT_FILE = "filtered-tweets.jsonl"  # Where we'll save matching Tweets

with open(INPUT_FILE_PATH, "r", encoding="utf-8") as infile, open(OUTPUT_FILE, "w", encoding="utf-8") as outfile:
    for line_number, line in enumerate(infile, 1):
        try:
            # Parse one Tweet from the line
            tweet = json.loads(line.strip())
            
            # Extract all hashtags from the Tweet (convert to lowercase for case-insensitive matching)
            tweet_hashtags = [tag["text"].lower() for tag in tweet.get("entities", {}).get("hashtags", [])]
            
            # Check if our target hashtag is present (we strip the # and lowercase it too)
            if TARGET_HASHTAG.lower().strip("#") in tweet_hashtags:
                # Write the matching Tweet to our output file
                json.dump(tweet, outfile, ensure_ascii=False)
                outfile.write("\n")  # Keep it as JSON Lines for consistency
                
        except json.JSONDecodeError:
            print(f"Skipping line {line_number}: Invalid JSON format")

Step 2: Make the Output More Readable

JSON Lines is great for processing, but not for reading. Let's convert the filtered Tweets to a CSV file—this will let you open it in Excel, Google Sheets, or even just a text editor easily:

import json
import csv

INPUT_FILTERED = "filtered-tweets.jsonl"  # Use the output from Step 1
OUTPUT_CSV = "readable-tweets.csv"

# Define which fields you want to keep (customize this list!)
DESIRED_FIELDS = [
    "created_at", 
    "id_str", 
    "text", 
    "user.screen_name",  # Nested field (user -> screen_name)
    "retweet_count", 
    "favorite_count"
]

with open(INPUT_FILTERED, "r", encoding="utf-8") as infile, open(OUTPUT_CSV, "w", newline="", encoding="utf-8") as outfile:
    writer = csv.DictWriter(outfile, fieldnames=DESIRED_FIELDS)
    writer.writeheader()  # Write column names first
    
    for line in infile:
        tweet = json.loads(line.strip())
        row = {}
        
        # Handle both top-level and nested fields
        for field in DESIRED_FIELDS:
            if "." in field:
                # Split nested fields (e.g., "user.screen_name" becomes ["user", "screen_name"])
                nested_parts = field.split(".")
                value = tweet
                for part in nested_parts:
                    value = value.get(part, "")  # Use empty string if field doesn't exist
                row[field] = value
            else:
                row[field] = tweet.get(field, "")
        
        writer.writerow(row)

Step 3: Sort the Tweets

If you want to sort by date (most recent first, for example), we'll first load the filtered Tweets into memory (this works if your filtered set is manageable; if it's still huge, we can use external tools, but let's start simple):

import json
from datetime import datetime

INPUT_FILTERED = "filtered-tweets.jsonl"
OUTPUT_SORTED = "sorted-tweets.jsonl"

# Load filtered Tweets and add a datetime field for sorting
tweets_list = []
with open(INPUT_FILTERED, "r", encoding="utf-8") as infile:
    for line in infile:
        tweet = json.loads(line.strip())
        # Convert the "created_at" string to a datetime object (Twitter's format: "Sun Oct 22 06:30:00 +0000 2017")
        tweet["created_datetime"] = datetime.strptime(tweet["created_at"], "%a %b %d %H:%M:%S %z %Y")
        tweets_list.append(tweet)

# Sort Tweets by date (reverse=True for newest first)
sorted_tweets = sorted(tweets_list, key=lambda x: x["created_datetime"], reverse=True)

# Write sorted Tweets to a new file (remove the temporary datetime field first)
with open(OUTPUT_SORTED, "w", encoding="utf-8") as outfile:
    for tweet in sorted_tweets:
        del tweet["created_datetime"]
        json.dump(tweet, outfile, ensure_ascii=False)
        outfile.write("\n")

Quick Tips for Newbies

  • File Paths: Make sure your input file path is correct—if the file is in the same folder as your script, you can just use the filename (e.g., archiveteam-twitter-stream-2017-01.json).
  • Memory Issues: If even the filtered file is too big to sort in memory, use command-line tools like sort (on Linux/macOS) or PowerShell's Sort-Object (Windows) to sort the JSON Lines file by the created_at field.
  • Jupyter Notebooks: If you want to use Jupyter, run the same scripts in cells—just avoid using read() on the entire file, stick to line-by-line loops.

内容的提问来源于stack exchange,提问作者pimwel

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 10:05:21