You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用MNB训练文本分类模型:如何移除指定类别及解决JSON读取问题

解决MNB文本分类任务中的数据集读取与类别过滤问题

1. 正确读取JSON格式的新闻数据集

该数据集的格式为每行一个独立JSON对象,而非标准的JSON数组结构,直接用json.load()读取会失败。可以通过逐行解析的方式读取:

import json
import pandas as pd

# 初始化空列表存储数据
data = []
# 逐行读取并解析JSON文件
with open('News_Category_Dataset_v3.json', 'r', encoding='utf-8') as f:
    for line in f:
        data.append(json.loads(line))

# 转换为DataFrame,方便后续数据处理
df = pd.DataFrame(data)

2. 移除不需要的类别数据

可以通过布尔索引或列表推导式过滤掉指定类别:

方法一:使用Pandas处理(推荐)

# 定义需要移除的类别
exclude_categories = ['US.NEWS', 'POLITICS']
# 过滤数据,保留不在排除列表中的类别
filtered_df = df[~df['category'].isin(exclude_categories)]
# 验证过滤结果
print(filtered_df['category'].unique())

方法二:纯列表处理(不依赖Pandas)

exclude_categories = ['US.NEWS', 'POLITICS']
# 筛选保留符合要求的数据项
filtered_data = [item for item in data if item['category'] not in exclude_categories]

数据集引用

  1. Misra, Rishabh and Prahal Arora. "Sarcasm Detection using Hybrid Neural Network." arXiv preprint arXiv:1908.07414 (2019)
  2. Misra, Rishabh and Jigyasa Grover. "Sculpting Data for ML: The first act of Machine Learning." ISBN 9798585463570 (2021)

内容的提问来源于stack exchange,提问作者Rijil Varghese

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 11:22:39