使用Python Tweepy采集Twitter数据:JSON转CSV遇无效JSON错误求解决方案
Hey there! Let's work through this issue you're having—converting Tweepy-scraped Twitter data from JSON to CSV keeps throwing an "invalid JSON" error. This usually boils down to how your JSON file is structured (super common with Tweepy outputs) or minor syntax glitches. Here's how to fix it step by step:
First: Diagnose the JSON Format Issue
Most of the time, the problem is that your JSON file uses JSON Lines format (one Tweet object per line) instead of a single valid JSON array. When you save Tweepy data by dumping each tweet individually with json.dump(), you end up with a file full of separate JSON objects, not one wrapped in [] with commas between them. This breaks standard JSON parsers.
Solution 1: Handle JSON Lines (Most Common Case)
If your tweets.json looks like this (one tweet per line):
{"id": 12345, "text": "Hello Twitter", ...} {"id": 67890, "text": "Another tweet", ...}
Use this code to read each line individually and write to CSV:
import json import csv # Open your JSON input and CSV output files with open('tweets.json', 'r', encoding='utf-8') as json_in, \ open('tweets_output.csv', 'w', newline='', encoding='utf-8') as csv_out: # Read the first tweet to get all field names (for CSV headers) first_tweet = json.loads(json_in.readline()) csv_headers = first_tweet.keys() # Set up the CSV writer writer = csv.DictWriter(csv_out, fieldnames=csv_headers) writer.writeheader() # Write the first tweet writer.writerow(first_tweet) # Process remaining lines, skipping any invalid JSON for line in json_in: try: tweet = json.loads(line.strip()) writer.writerow(tweet) except json.JSONDecodeError as e: print(f"Skipping invalid line: {str(e)}") continue
Solution 2: Fix a Malformed JSON Array
If you intended to save your data as a single JSON array but messed up the syntax (missing [/], extra commas, etc.), first repair the JSON:
- Run this command in your terminal to validate and format the file (it'll tell you where errors are):
python -m json.tool tweets.json fixed_tweets.json - Manually fix any syntax issues the tool flags (like a missing closing bracket or extra comma at the end of the array).
Once your JSON is a valid array, use this code to convert it:
import json import csv with open('fixed_tweets.json', 'r', encoding='utf-8') as json_in, \ open('tweets_output.csv', 'w', newline='', encoding='utf-8') as csv_out: tweets = json.load(json_in) if not tweets: print("No tweets found in the JSON file.") exit() csv_headers = tweets[0].keys() writer = csv.DictWriter(csv_out, fieldnames=csv_headers) writer.writeheader() for tweet in tweets: writer.writerow(tweet)
Bonus: Flatten Nested Tweet Fields
Tweepy returns nested data (like user objects, entities, etc.), which will show up as raw strings in your CSV. If you want to expand these into separate columns, add a helper function to flatten the tweet data:
def flatten_tweet(tweet): flattened = tweet.copy() # Expand user details into separate columns if 'user' in flattened: flattened['user_name'] = flattened['user']['name'] flattened['user_screen_name'] = flattened['user']['screen_name'] flattened['user_followers_count'] = flattened['user']['followers_count'] del flattened['user'] # Add more fields to flatten here (e.g., entities, retweeted_status) return flattened # Use it in your loop like this: writer.writerow(flatten_tweet(tweet))
Key Tips to Avoid Future Issues
- Save correctly from Tweepy: If you want a valid JSON array from the start, collect all tweets in a list first, then dump the entire list:
tweets_list = [] # ... your Tweepy scraping code ... tweets_list.append(tweet._json) # Use _json to get the raw JSON object with open('tweets.json', 'w', encoding='utf-8') as f: json.dump(tweets_list, f, indent=2) - Use UTF-8 encoding: Always specify
encoding='utf-8'when reading/writing files to avoid character encoding errors. - Add error handling: The
try-exceptblocks in Solution 1 will prevent your script from crashing if it hits a bad line.
内容的提问来源于stack exchange,提问作者Akshma Mittal

