Python转换Hydrated Tweets JSONL到CSV时遇JSONDecodeError报错求助
JSONL转CSV报错排查与解决方法
我用以下代码将Hydrated Tweets的JSONL文件转换为CSV格式,此前处理多个数据集均正常,但现在运行报错:
import pandas as pd import json # Json file name with open('climate.jsonl') as f: lines = f.read().splitlines() print('jline opened') df_inter = pd.DataFrame(lines) df_inter.columns = ['json_element'] print('df_inter.columns') df_inter['json_element'].apply(json.loads) print('df_inter') Twitter_Dataset = pd.json_normalize(df_inter['json_element'].apply(json.loads)) # Output CSV file Twitter_Dataset.to_csv('Sandy_tweets_unfiltered.csv') print('Unfiltered tweets have successfuly saved as CSV')
报错信息:
Traceback (most recent call last): File "2-Json_to_CSV.py", line 15, in <module> df_inter['json_element'].apply(json.loads) File "/share/apps/miniconda3/installed_1/envs/sgupta5/lib/python3.7/site-packages/pandas/core/series.py", line 4200, in apply mapped = lib.map_infer(values, f, convert=convert_dtype) File "pandas/_libs/lib.pyx", line 2402, in pandas._libs.lib.map_infer File "/share/apps/miniconda3/installed_1/envs/sgupta5/lib/python3.7/json/__init__.py", line 348, in loads return _default_decoder.decode(s) File "/share/apps/miniconda3/installed_1/envs/sgupta5/lib/python3.7/json/decoder.py", line 340, in decode raise JSONDecodeError("Extra data", s, end) json.decoder.JSONDecodeError: Extra data: line 1 column 13 (char 12) srun: error: compute-20-9: task 0: Exited with exit code 1
错误原因
JSONDecodeError: Extra data 说明目标JSONL文件中存在格式不合法的行,常见情况:
- 某一行包含多个独立的JSON对象(比如一行打包了两条Tweet数据)
- 某一行JSON语法错误(引号不闭合、多余逗号、格式混乱)
- 文件中存在空行或非JSON格式的无效内容
解决步骤
1. 定位错误行
先运行这段代码,找出具体出问题的行号和内容:
import json error_lines = [] with open('climate.jsonl') as f: for line_num, line in enumerate(f, 1): line = line.strip() if not line: continue try: json.loads(line) except json.JSONDecodeError as e: error_lines.append((line_num, line[:50] + "...", str(e))) if error_lines: print("发现错误行:") for num, content, err in error_lines: print(f"行号: {num}, 内容片段: {content}, 错误: {err}") else: print("所有行JSON格式均合法")
2. 处理错误行
根据定位结果选择对应方案:
- 空行问题:在原代码读取时直接过滤空行:
with open('climate.jsonl') as f: lines = [line.strip() for line in f if line.strip()] - 一行多JSON:根据具体格式拆分该行(比如用逗号分割后逐个解析),或直接跳过该行
- 语法错误行:少量错误可直接跳过;大量错误则需检查数据源是否损坏
3. 优化后的完整代码
加入错误处理,避免程序崩溃,同时过滤无效数据:
import pandas as pd import json def safe_json_parse(line): try: return json.loads(line) except json.JSONDecodeError: return None # 读取并过滤空行 with open('climate.jsonl') as f: lines = [line.strip() for line in f if line.strip()] # 解析JSON并过滤无效行 df = pd.DataFrame(lines, columns=['raw_json']) df['parsed_json'] = df['raw_json'].apply(safe_json_parse) df_valid = df.dropna(subset=['parsed_json']) # 转换为CSV并保存 if not df_valid.empty: twitter_data = pd.json_normalize(df_valid['parsed_json']) twitter_data.to_csv('Sandy_tweets_unfiltered.csv', index=False) print(f"成功保存{len(twitter_data)}条有效数据,跳过{len(df)-len(df_valid)}条错误行") else: print("无有效数据可保存")
内容的提问来源于stack exchange,提问作者Kasperr
相关产品推荐
相关产品推荐

