You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

合并多个JSONL文件并使用Python读取时出现JSONDecodeError求助

解决JSONL文件合并与读取的JSONDecodeError问题

咱们先定位核心问题:你当前的合并代码并没有实际将数据写入到output.jsonl文件中(注释掉的写入代码写法也不对),而且就算你尝试把tweets列表直接转成JSON写入,得到的会是一个JSON数组,完全不符合JSONL(每行一个独立JSON对象)的格式要求,这就是读取时解码报错的根源。

修正后的JSONL合并代码

正确的做法是保持JSONL的格式特性:逐个读取源文件的每一行,确保目标文件的每行都是一个独立的JSON对象,代码如下:

import glob
import json

# 匹配所有源JSONL文件
json_files = glob.glob('C:\\Users\\arun\\Desktop\\Tweets\\*.jsonl')
output_path = 'C:\\Users\\arun\\Desktop\\Tweets\\output.jsonl'

# 打开目标文件,逐行写入有效JSON对象
with open(output_path, 'w', encoding='utf-8') as outfile:
    for file_path in json_files:
        with open(file_path, 'r', encoding='utf-8') as infile:
            for line in infile:
                # 跳过空行,避免无效内容
                clean_line = line.strip()
                if not clean_line:
                    continue
                # 加载后重新序列化,保证格式统一
                tweet = json.loads(clean_line)
                json.dump(tweet, outfile, ensure_ascii=False)
                outfile.write('\n')  # 强制换行,维持JSONL格式

修正后的读取代码

读取时建议加入异常捕获,避免单个无效行导致整个程序崩溃,同时完善标签统计逻辑:

from collections import Counter
import json

def get_hashtags(tweet):
    entities = tweet.get('entities', {})
    hashtags = entities.get('hashtags', [])
    return [tag['text'].lower() for tag in hashtags]

fname = "C:\\Users\\arun\\Desktop\\Tweets\\output.jsonl"
hashtags_counter = Counter()

with open(fname, 'r', encoding='utf-8') as f:
    for line_num, line in enumerate(f, 1):
        clean_line = line.strip()
        if not clean_line:
            continue
        try:
            tweet = json.loads(clean_line)
            # 提取标签并更新统计
            tweet_hashtags = get_hashtags(tweet)
            hashtags_counter.update(tweet_hashtags)
        except json.JSONDecodeError as e:
            print(f"第{line_num}行解析失败: {e}")
            continue

# 打印Top10热门标签示例
print("热门标签统计结果:")
for tag, count in hashtags_counter.most_common(10):
    print(f"{tag}: {count}")

关键注意事项

  • JSONL格式规则:必须保证每行是一个独立的JSON对象,不能是数组或多对象用逗号拼接,否则json.loads(line)无法解析单行内容。
  • 编码一致性:读写文件时指定encoding='utf-8',避免特殊字符或非英文内容出现乱码或解析失败。
  • 容错处理:添加异常捕获后,程序可以跳过无效行,不会因为个别错误中断整个读取流程。

内容的提问来源于stack exchange,提问作者Balaji Venky

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 09:36:55