合并多个JSONL文件并使用Python读取时出现JSONDecodeError求助
解决JSONL文件合并与读取的JSONDecodeError问题
咱们先定位核心问题:你当前的合并代码并没有实际将数据写入到output.jsonl文件中(注释掉的写入代码写法也不对),而且就算你尝试把tweets列表直接转成JSON写入,得到的会是一个JSON数组,完全不符合JSONL(每行一个独立JSON对象)的格式要求,这就是读取时解码报错的根源。
修正后的JSONL合并代码
正确的做法是保持JSONL的格式特性:逐个读取源文件的每一行,确保目标文件的每行都是一个独立的JSON对象,代码如下:
import glob import json # 匹配所有源JSONL文件 json_files = glob.glob('C:\\Users\\arun\\Desktop\\Tweets\\*.jsonl') output_path = 'C:\\Users\\arun\\Desktop\\Tweets\\output.jsonl' # 打开目标文件,逐行写入有效JSON对象 with open(output_path, 'w', encoding='utf-8') as outfile: for file_path in json_files: with open(file_path, 'r', encoding='utf-8') as infile: for line in infile: # 跳过空行,避免无效内容 clean_line = line.strip() if not clean_line: continue # 加载后重新序列化,保证格式统一 tweet = json.loads(clean_line) json.dump(tweet, outfile, ensure_ascii=False) outfile.write('\n') # 强制换行,维持JSONL格式
修正后的读取代码
读取时建议加入异常捕获,避免单个无效行导致整个程序崩溃,同时完善标签统计逻辑:
from collections import Counter import json def get_hashtags(tweet): entities = tweet.get('entities', {}) hashtags = entities.get('hashtags', []) return [tag['text'].lower() for tag in hashtags] fname = "C:\\Users\\arun\\Desktop\\Tweets\\output.jsonl" hashtags_counter = Counter() with open(fname, 'r', encoding='utf-8') as f: for line_num, line in enumerate(f, 1): clean_line = line.strip() if not clean_line: continue try: tweet = json.loads(clean_line) # 提取标签并更新统计 tweet_hashtags = get_hashtags(tweet) hashtags_counter.update(tweet_hashtags) except json.JSONDecodeError as e: print(f"第{line_num}行解析失败: {e}") continue # 打印Top10热门标签示例 print("热门标签统计结果:") for tag, count in hashtags_counter.most_common(10): print(f"{tag}: {count}")
关键注意事项
- JSONL格式规则:必须保证每行是一个独立的JSON对象,不能是数组或多对象用逗号拼接,否则
json.loads(line)无法解析单行内容。 - 编码一致性:读写文件时指定
encoding='utf-8',避免特殊字符或非英文内容出现乱码或解析失败。 - 容错处理:添加异常捕获后,程序可以跳过无效行,不会因为个别错误中断整个读取流程。
内容的提问来源于stack exchange,提问作者Balaji Venky
相关产品推荐
相关产品推荐

