You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python转换Hydrated Tweets JSONL到CSV时遇JSONDecodeError报错求助

JSONL转CSV报错排查与解决方法

我用以下代码将Hydrated Tweets的JSONL文件转换为CSV格式,此前处理多个数据集均正常,但现在运行报错:

import pandas as pd
import json

# Json file name
with open('climate.jsonl') as f:
    lines = f.read().splitlines()
print('jline opened')
df_inter = pd.DataFrame(lines)
df_inter.columns = ['json_element']

print('df_inter.columns')

df_inter['json_element'].apply(json.loads)

print('df_inter')

Twitter_Dataset = pd.json_normalize(df_inter['json_element'].apply(json.loads))

# Output CSV file
Twitter_Dataset.to_csv('Sandy_tweets_unfiltered.csv')
print('Unfiltered tweets have successfuly saved as CSV')

报错信息:

Traceback (most recent call last):
  File "2-Json_to_CSV.py", line 15, in <module>
    df_inter['json_element'].apply(json.loads)
  File "/share/apps/miniconda3/installed_1/envs/sgupta5/lib/python3.7/site-packages/pandas/core/series.py", line 4200, in apply
    mapped = lib.map_infer(values, f, convert=convert_dtype)
  File "pandas/_libs/lib.pyx", line 2402, in pandas._libs.lib.map_infer
  File "/share/apps/miniconda3/installed_1/envs/sgupta5/lib/python3.7/json/__init__.py", line 348, in loads
    return _default_decoder.decode(s)
  File "/share/apps/miniconda3/installed_1/envs/sgupta5/lib/python3.7/json/decoder.py", line 340, in decode
    raise JSONDecodeError("Extra data", s, end)
json.decoder.JSONDecodeError: Extra data: line 1 column 13 (char 12)
srun: error: compute-20-9: task 0: Exited with exit code 1

错误原因

JSONDecodeError: Extra data 说明目标JSONL文件中存在格式不合法的行,常见情况:

  • 某一行包含多个独立的JSON对象(比如一行打包了两条Tweet数据)
  • 某一行JSON语法错误(引号不闭合、多余逗号、格式混乱)
  • 文件中存在空行或非JSON格式的无效内容

解决步骤

1. 定位错误行

先运行这段代码,找出具体出问题的行号和内容:

import json

error_lines = []
with open('climate.jsonl') as f:
    for line_num, line in enumerate(f, 1):
        line = line.strip()
        if not line:
            continue
        try:
            json.loads(line)
        except json.JSONDecodeError as e:
            error_lines.append((line_num, line[:50] + "...", str(e)))

if error_lines:
    print("发现错误行:")
    for num, content, err in error_lines:
        print(f"行号: {num}, 内容片段: {content}, 错误: {err}")
else:
    print("所有行JSON格式均合法")

2. 处理错误行

根据定位结果选择对应方案:

  • 空行问题:在原代码读取时直接过滤空行:
    with open('climate.jsonl') as f:
        lines = [line.strip() for line in f if line.strip()]
    
  • 一行多JSON:根据具体格式拆分该行(比如用逗号分割后逐个解析),或直接跳过该行
  • 语法错误行:少量错误可直接跳过;大量错误则需检查数据源是否损坏

3. 优化后的完整代码

加入错误处理,避免程序崩溃,同时过滤无效数据:

import pandas as pd
import json

def safe_json_parse(line):
    try:
        return json.loads(line)
    except json.JSONDecodeError:
        return None

# 读取并过滤空行
with open('climate.jsonl') as f:
    lines = [line.strip() for line in f if line.strip()]

# 解析JSON并过滤无效行
df = pd.DataFrame(lines, columns=['raw_json'])
df['parsed_json'] = df['raw_json'].apply(safe_json_parse)
df_valid = df.dropna(subset=['parsed_json'])

# 转换为CSV并保存
if not df_valid.empty:
    twitter_data = pd.json_normalize(df_valid['parsed_json'])
    twitter_data.to_csv('Sandy_tweets_unfiltered.csv', index=False)
    print(f"成功保存{len(twitter_data)}条有效数据,跳过{len(df)-len(df_valid)}条错误行")
else:
    print("无有效数据可保存")

内容的提问来源于stack exchange,提问作者Kasperr

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.21 08:48:21