新手求助:无法处理Archive.org的Twitter JSON数据集
Hey there! No worries at all—everyone starts somewhere, and handling massive Twitter JSON datasets can be tricky even for seasoned developers. Let's walk through how to fix your issues step by step.
First, let's clarify what you're working with: those Archive.org Twitter files are JSON Lines (.jsonl) files, meaning each line is a separate JSON object (one Tweet per line). That's why trying to open the whole file at once crashes your tools—it's way too big to fit in memory. We'll use line-by-line processing to get around that.
Step 1: Read the File & Filter by Hashtag
Instead of loading the entire file into memory, we'll read one line at a time, parse the Tweet, and check if it contains your target hashtag. Here's a simple Python script to do that:
import json from datetime import datetime # Customize these variables to match your needs TARGET_HASHTAG = "#YourHashtagHere" # Replace with your desired hashtag INPUT_FILE_PATH = "path/to/your/twitter-file.json" # Update this path OUTPUT_FILE = "filtered-tweets.jsonl" # Where we'll save matching Tweets with open(INPUT_FILE_PATH, "r", encoding="utf-8") as infile, open(OUTPUT_FILE, "w", encoding="utf-8") as outfile: for line_number, line in enumerate(infile, 1): try: # Parse one Tweet from the line tweet = json.loads(line.strip()) # Extract all hashtags from the Tweet (convert to lowercase for case-insensitive matching) tweet_hashtags = [tag["text"].lower() for tag in tweet.get("entities", {}).get("hashtags", [])] # Check if our target hashtag is present (we strip the # and lowercase it too) if TARGET_HASHTAG.lower().strip("#") in tweet_hashtags: # Write the matching Tweet to our output file json.dump(tweet, outfile, ensure_ascii=False) outfile.write("\n") # Keep it as JSON Lines for consistency except json.JSONDecodeError: print(f"Skipping line {line_number}: Invalid JSON format")
Step 2: Make the Output More Readable
JSON Lines is great for processing, but not for reading. Let's convert the filtered Tweets to a CSV file—this will let you open it in Excel, Google Sheets, or even just a text editor easily:
import json import csv INPUT_FILTERED = "filtered-tweets.jsonl" # Use the output from Step 1 OUTPUT_CSV = "readable-tweets.csv" # Define which fields you want to keep (customize this list!) DESIRED_FIELDS = [ "created_at", "id_str", "text", "user.screen_name", # Nested field (user -> screen_name) "retweet_count", "favorite_count" ] with open(INPUT_FILTERED, "r", encoding="utf-8") as infile, open(OUTPUT_CSV, "w", newline="", encoding="utf-8") as outfile: writer = csv.DictWriter(outfile, fieldnames=DESIRED_FIELDS) writer.writeheader() # Write column names first for line in infile: tweet = json.loads(line.strip()) row = {} # Handle both top-level and nested fields for field in DESIRED_FIELDS: if "." in field: # Split nested fields (e.g., "user.screen_name" becomes ["user", "screen_name"]) nested_parts = field.split(".") value = tweet for part in nested_parts: value = value.get(part, "") # Use empty string if field doesn't exist row[field] = value else: row[field] = tweet.get(field, "") writer.writerow(row)
Step 3: Sort the Tweets
If you want to sort by date (most recent first, for example), we'll first load the filtered Tweets into memory (this works if your filtered set is manageable; if it's still huge, we can use external tools, but let's start simple):
import json from datetime import datetime INPUT_FILTERED = "filtered-tweets.jsonl" OUTPUT_SORTED = "sorted-tweets.jsonl" # Load filtered Tweets and add a datetime field for sorting tweets_list = [] with open(INPUT_FILTERED, "r", encoding="utf-8") as infile: for line in infile: tweet = json.loads(line.strip()) # Convert the "created_at" string to a datetime object (Twitter's format: "Sun Oct 22 06:30:00 +0000 2017") tweet["created_datetime"] = datetime.strptime(tweet["created_at"], "%a %b %d %H:%M:%S %z %Y") tweets_list.append(tweet) # Sort Tweets by date (reverse=True for newest first) sorted_tweets = sorted(tweets_list, key=lambda x: x["created_datetime"], reverse=True) # Write sorted Tweets to a new file (remove the temporary datetime field first) with open(OUTPUT_SORTED, "w", encoding="utf-8") as outfile: for tweet in sorted_tweets: del tweet["created_datetime"] json.dump(tweet, outfile, ensure_ascii=False) outfile.write("\n")
Quick Tips for Newbies
- File Paths: Make sure your input file path is correct—if the file is in the same folder as your script, you can just use the filename (e.g.,
archiveteam-twitter-stream-2017-01.json). - Memory Issues: If even the filtered file is too big to sort in memory, use command-line tools like
sort(on Linux/macOS) or PowerShell'sSort-Object(Windows) to sort the JSON Lines file by thecreated_atfield. - Jupyter Notebooks: If you want to use Jupyter, run the same scripts in cells—just avoid using
read()on the entire file, stick to line-by-line loops.
内容的提问来源于stack exchange,提问作者pimwel

