You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用NLTK处理OSHA文本数据:分类及筛选返回零行问题求助

解决OSHA文本数据处理的两个问题

问题2:筛选指定类别返回零行的解决方法

核心原因是你编写的筛选字符串与Class列的实际值不匹配,比如大小写差异(employee vs Employee)、前后空格或格式不一致。

解决步骤:

  • 先确认Class列的所有唯一值,获取准确的类别名称:
print(news2['Class'].unique())
  • 根据输出的实际值调整筛选条件,比如实际值是"09 Employee Injures Hand",就把筛选语句中的对应字符串修正为准确值;
  • 或者先统一清理Class列的格式,避免格式干扰:
# 去除前后空格,统一首字母大写(可按需调整规则)
news2['Class'] = news2['Class'].str.strip().str.title()
# 用isin简化多条件筛选
subnews2 = news2[news2['Class'].isin(["09 Employee Injures Hand", "29 Employees Burned", "3 Employees Killed"])]
print(subnews2.shape)

问题1:无法将信息分类到对应类别的解决方法

要实现文本到指定类别的分类,需要完整的NLTK文本分类流程,步骤如下:

  1. 文本预处理:清洗文本,提取有效特征
from nltk.corpus import stopwords
from nltk.tokenize import word_tokenize
from nltk.stem import WordNetLemmatizer
nltk.download('wordnet')

stop_words = set(stopwords.words('english'))
lemmatizer = WordNetLemmatizer()

def preprocess_text(text):
    # 分词并转小写
    tokens = word_tokenize(text.lower())
    # 过滤停用词、非字母字符,做词形还原
    filtered_tokens = [lemmatizer.lemmatize(token) for token in tokens if token.isalpha() and token not in stop_words]
    return ' '.join(filtered_tokens)

# 应用预处理到文本列
news2['Cleaned_Text'] = news2['Text'].apply(preprocess_text)
  1. 提取数值特征:用TF-IDF将文本转为模型可识别的数值格式
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.model_selection import train_test_split

tfidf = TfidfVectorizer(max_features=1000)
X = tfidf.fit_transform(news2['Cleaned_Text']).toarray()
y = news2['Class']

# 划分训练集和测试集
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
  1. 训练并评估分类模型:以朴素贝叶斯为例
from sklearn.naive_bayes import MultinomialNB
from sklearn.metrics import classification_report

# 训练模型
model = MultinomialNB()
model.fit(X_train, y_train)

# 评估模型效果
y_pred = model.predict(X_test)
print(classification_report(y_test, y_pred))
  1. 使用模型分类新文本
def classify_text(text):
    cleaned_text = preprocess_text(text)
    text_features = tfidf.transform([cleaned_text]).toarray()
    return model.predict(text_features)[0]

# 示例:对新文本分类
sample_text = "Worker suffered hand injury while operating industrial equipment"
print(classify_text(sample_text))

内容的提问来源于stack exchange,提问作者Rochhh999

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 14:05:31