You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python Tweepy采集Twitter数据:JSON转CSV遇无效JSON错误求解决方案

Fixing Invalid JSON Error When Converting Tweepy Data to CSV

Hey there! Let's work through this issue you're having—converting Tweepy-scraped Twitter data from JSON to CSV keeps throwing an "invalid JSON" error. This usually boils down to how your JSON file is structured (super common with Tweepy outputs) or minor syntax glitches. Here's how to fix it step by step:

First: Diagnose the JSON Format Issue

Most of the time, the problem is that your JSON file uses JSON Lines format (one Tweet object per line) instead of a single valid JSON array. When you save Tweepy data by dumping each tweet individually with json.dump(), you end up with a file full of separate JSON objects, not one wrapped in [] with commas between them. This breaks standard JSON parsers.

Solution 1: Handle JSON Lines (Most Common Case)

If your tweets.json looks like this (one tweet per line):

{"id": 12345, "text": "Hello Twitter", ...}
{"id": 67890, "text": "Another tweet", ...}

Use this code to read each line individually and write to CSV:

import json
import csv

# Open your JSON input and CSV output files
with open('tweets.json', 'r', encoding='utf-8') as json_in, \
     open('tweets_output.csv', 'w', newline='', encoding='utf-8') as csv_out:
    
    # Read the first tweet to get all field names (for CSV headers)
    first_tweet = json.loads(json_in.readline())
    csv_headers = first_tweet.keys()
    
    # Set up the CSV writer
    writer = csv.DictWriter(csv_out, fieldnames=csv_headers)
    writer.writeheader()
    
    # Write the first tweet
    writer.writerow(first_tweet)
    
    # Process remaining lines, skipping any invalid JSON
    for line in json_in:
        try:
            tweet = json.loads(line.strip())
            writer.writerow(tweet)
        except json.JSONDecodeError as e:
            print(f"Skipping invalid line: {str(e)}")
            continue

Solution 2: Fix a Malformed JSON Array

If you intended to save your data as a single JSON array but messed up the syntax (missing [/], extra commas, etc.), first repair the JSON:

  1. Run this command in your terminal to validate and format the file (it'll tell you where errors are):
    python -m json.tool tweets.json fixed_tweets.json
    
  2. Manually fix any syntax issues the tool flags (like a missing closing bracket or extra comma at the end of the array).

Once your JSON is a valid array, use this code to convert it:

import json
import csv

with open('fixed_tweets.json', 'r', encoding='utf-8') as json_in, \
     open('tweets_output.csv', 'w', newline='', encoding='utf-8') as csv_out:
    
    tweets = json.load(json_in)
    if not tweets:
        print("No tweets found in the JSON file.")
        exit()
    
    csv_headers = tweets[0].keys()
    writer = csv.DictWriter(csv_out, fieldnames=csv_headers)
    writer.writeheader()
    
    for tweet in tweets:
        writer.writerow(tweet)

Bonus: Flatten Nested Tweet Fields

Tweepy returns nested data (like user objects, entities, etc.), which will show up as raw strings in your CSV. If you want to expand these into separate columns, add a helper function to flatten the tweet data:

def flatten_tweet(tweet):
    flattened = tweet.copy()
    
    # Expand user details into separate columns
    if 'user' in flattened:
        flattened['user_name'] = flattened['user']['name']
        flattened['user_screen_name'] = flattened['user']['screen_name']
        flattened['user_followers_count'] = flattened['user']['followers_count']
        del flattened['user']
    
    # Add more fields to flatten here (e.g., entities, retweeted_status)
    return flattened

# Use it in your loop like this:
writer.writerow(flatten_tweet(tweet))

Key Tips to Avoid Future Issues

  • Save correctly from Tweepy: If you want a valid JSON array from the start, collect all tweets in a list first, then dump the entire list:
    tweets_list = []
    # ... your Tweepy scraping code ...
    tweets_list.append(tweet._json)  # Use _json to get the raw JSON object
    
    with open('tweets.json', 'w', encoding='utf-8') as f:
        json.dump(tweets_list, f, indent=2)
    
  • Use UTF-8 encoding: Always specify encoding='utf-8' when reading/writing files to avoid character encoding errors.
  • Add error handling: The try-except blocks in Solution 1 will prevent your script from crashing if it hits a bad line.

内容的提问来源于stack exchange,提问作者Akshma Mittal

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 08:44:38