You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何优化DataFrame文本情感极性计算方法以缩减运行耗时

优化方案

1 替换词表为集合,消除O(n)查询开销

原有代码使用列表做in查询,每次匹配需要遍历最多2000个元素,替换为集合后查询复杂度降到O(1),是成本最低收益最高的优化,仅这一步即可提速5-10倍。

# 提前把词表转成集合,只需执行一次
positive_set = set(positive_words)
negative_set = set(negative_words)

def get_sentiment_score(text):
    words = text.split()
    positive = 0
    negative = 0
    for word in words:
        if word in positive_set:
            positive += 1
        elif word in negative_set:
            negative += 1
    score = positive - negative
    if score == 0:
        return "UNCERTAIN"
    return "POSITIVE" if score > 0 else "NEGATIVE"

# 仍保留apply的前提下,速度已经有大幅提升
df["sentiment_polarity"] = df["text"].apply(lambda row: get_sentiment_score(str(row)))

2 放弃逐行apply,使用向量化批量计算(推荐)

pandas原生apply是Python层的逐行迭代,27万行的循环开销很大,使用sklearn的CountVectorizer做批量词频统计,完全用numpy矩阵运算代替Python循环,总耗时可降到10秒以内,远低于100秒的要求。

注意设置token_pattern=r'\S+'是为了和原有split()的分词逻辑完全对齐,不会改变原有分类结果。

from sklearn.feature_extraction.text import CountVectorizer
import numpy as np

# 提前合并正负词表,指定Vectorizer的词表,避免统计无关词
all_vocab = positive_words + negative_words
vectorizer = CountVectorizer(vocabulary=all_vocab, lowercase=False, token_pattern=r'\S+')
# 批量统计所有文本的词频,得到的是27万行 * 4000+列的稀疏矩阵
word_counts = vectorizer.transform(df["text"].astype(str))

# 计算每个样本的正负词总数量
pos_count = word_counts[:, :len(positive_words)].sum(axis=1).A1
neg_count = word_counts[:, len(positive_words):].sum(axis=1).A1
score = pos_count - neg_count

# 批量生成分类结果
conditions = [score > 0, score < 0, score == 0]
choices = ["POSITIVE", "NEGATIVE", "UNCERTAIN"]
df["sentiment_polarity"] = np.select(conditions, choices)

3 保留自定义逻辑的可选方案:用numba编译加速

如果后续需要在函数中增加更复杂的自定义规则,不想放弃逐行处理逻辑,可以用numba的JIT编译把函数转成机器码,循环速度和C语言相当,提速效果也能达到10倍以上。

from numba import njit

# 提前把词转成numba支持的字典映射
positive_dict = {word:1 for word in positive_words}
negative_dict = {word:1 for word in negative_words}

@njit
def get_sentiment_score_numba(text):
    words = text.split()
    positive = 0
    negative = 0
    for word in words:
        if word in positive_dict:
            positive += 1
        elif word in negative_dict:
            negative += 1
    score = positive - negative
    if score == 0:
        return "UNCERTAIN"
    return "POSITIVE" if score > 0 else "NEGATIVE"

df["sentiment_polarity"] = df["text"].apply(lambda row: get_sentiment_score_numba(str(row)))

内容的提问来源于stack exchange,提问作者Piyush Jain

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.05 14:06:04