如何优化DataFrame文本情感极性计算方法以缩减运行耗时
优化方案
1 替换词表为集合,消除O(n)查询开销
原有代码使用列表做in查询,每次匹配需要遍历最多2000个元素,替换为集合后查询复杂度降到O(1),是成本最低收益最高的优化,仅这一步即可提速5-10倍。
# 提前把词表转成集合,只需执行一次 positive_set = set(positive_words) negative_set = set(negative_words) def get_sentiment_score(text): words = text.split() positive = 0 negative = 0 for word in words: if word in positive_set: positive += 1 elif word in negative_set: negative += 1 score = positive - negative if score == 0: return "UNCERTAIN" return "POSITIVE" if score > 0 else "NEGATIVE" # 仍保留apply的前提下,速度已经有大幅提升 df["sentiment_polarity"] = df["text"].apply(lambda row: get_sentiment_score(str(row)))
2 放弃逐行apply,使用向量化批量计算(推荐)
pandas原生apply是Python层的逐行迭代,27万行的循环开销很大,使用sklearn的CountVectorizer做批量词频统计,完全用numpy矩阵运算代替Python循环,总耗时可降到10秒以内,远低于100秒的要求。
注意设置
token_pattern=r'\S+'是为了和原有split()的分词逻辑完全对齐,不会改变原有分类结果。
from sklearn.feature_extraction.text import CountVectorizer import numpy as np # 提前合并正负词表,指定Vectorizer的词表,避免统计无关词 all_vocab = positive_words + negative_words vectorizer = CountVectorizer(vocabulary=all_vocab, lowercase=False, token_pattern=r'\S+') # 批量统计所有文本的词频,得到的是27万行 * 4000+列的稀疏矩阵 word_counts = vectorizer.transform(df["text"].astype(str)) # 计算每个样本的正负词总数量 pos_count = word_counts[:, :len(positive_words)].sum(axis=1).A1 neg_count = word_counts[:, len(positive_words):].sum(axis=1).A1 score = pos_count - neg_count # 批量生成分类结果 conditions = [score > 0, score < 0, score == 0] choices = ["POSITIVE", "NEGATIVE", "UNCERTAIN"] df["sentiment_polarity"] = np.select(conditions, choices)
3 保留自定义逻辑的可选方案:用numba编译加速
如果后续需要在函数中增加更复杂的自定义规则,不想放弃逐行处理逻辑,可以用numba的JIT编译把函数转成机器码,循环速度和C语言相当,提速效果也能达到10倍以上。
from numba import njit # 提前把词转成numba支持的字典映射 positive_dict = {word:1 for word in positive_words} negative_dict = {word:1 for word in negative_words} @njit def get_sentiment_score_numba(text): words = text.split() positive = 0 negative = 0 for word in words: if word in positive_dict: positive += 1 elif word in negative_dict: negative += 1 score = positive - negative if score == 0: return "UNCERTAIN" return "POSITIVE" if score > 0 else "NEGATIVE" df["sentiment_polarity"] = df["text"].apply(lambda row: get_sentiment_score_numba(str(row)))
内容的提问来源于stack exchange,提问作者Piyush Jain
相关产品推荐
相关产品推荐

