如何在Python中清洗推文并提取JSON文件的full_text字段?
full_text from Tweet JSON File Hey there! Let's walk through how to clean your tweet data by extracting the full_text field from your JSON file. This is a common task when working with Twitter datasets, and Python's built-in json module makes it straightforward.
Step-by-Step Solution
First, we'll load the JSON file, parse its contents, then pull out the full_text values while handling edge cases (like tweets missing the field).
Here's the complete code:
import json # Load the raw tweets JSON file with open('raw_tweets.json', 'r', encoding='utf-8') as input_file: tweets = json.load(input_file) # Extract and collect the full_text fields cleaned_tweets = [] for tweet in tweets: # Check if the full_text field exists to avoid KeyError if 'full_text' in tweet: cleaned_tweets.append(tweet['full_text']) else: # Optional: Log tweets that don't have the field print(f"Skipping tweet (no full_text): {tweet.get('id_str', 'unknown ID')}") # Example: Print the cleaned tweets (matches your desired output format) for text in cleaned_tweets: print(text) # Optional: Save cleaned tweets to a text file with open('cleaned_tweets.txt', 'w', encoding='utf-8') as output_file: for text in cleaned_tweets: output_file.write(f"{text}\n")
Key Notes
- File Handling: Using
withensures the file is properly closed after reading/writing, which is Python best practice. - Error Prevention: We check if
full_textexists in each tweet to avoid crashing if some entries don't have the field. - Encoding: Using
encoding='utf-8'ensures we handle special characters (like emojis or ampersands in your example) correctly.
When you run this code, you'll get output exactly like the sample you provided:
#Deathstroke 31 @DCComics • super airDrop opening & it only gets better from there • it’s not just #BatmanMammaMia, folks! #SladeWilson in dept to #Mento & #BruceWayne methodology a bit more cosmopolitan https://t.co/jWUGBn4Fqm
内容的提问来源于stack exchange,提问作者bybu

