You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何提升基于Sentiment140数据集的情感分析Python算法运行速度

核心性能瓶颈分析

原代码的时间复杂度为O(未知词数量 * 总推文数),在百万级未知词的场景下属于指数级耗时,主要冗余点有三个:

  • 未知词列表存在大量重复值:同个词多次出现在不同推文中时,会被重复加入words_without_embs,导致重复计算
  • 每次找未知词的出现位置都全量遍历所有推文,没有提前做索引映射
  • 循环内冗余判断、裸异常捕获、重复向量计算都会拖慢速度

具体优化方案

第一步:先对未知词去重,同时预构建词的倒排索引

提前统计每个未知词在所有推文中的出现位置,一次遍历推文就能完成,后续不用重复遍历:

from collections import defaultdict

# 1. 去重未知词,同时构建词 -> 出现位置(推文下标, 词在推文中的下标)的倒排索引
word_positions = defaultdict(list)
unique_oov_words = set() # OOV即未登录词/未知词
for tweet_idx, tweet in enumerate(tweet_tokens):
    for word_idx, word in enumerate(tweet):
        if word not in word2vec:
            unique_oov_words.add(word)
            word_positions[word].append((tweet_idx, word_idx))

# 此时unique_oov_words就是去重后的未知词,数量会远小于1058532

第二步:优化向量计算逻辑,去掉冗余循环和判断

直接根据倒排索引取上下文,避免全量遍历推文,同时显式判断下标边界,替换低效的裸异常捕获:

vectors = {}
emb_dim = word2vec.vector_size # 提前拿向量维度,避免重复读取

for word in unique_oov_words:
    context_vecs = []
    # 直接取该词所有出现位置,不用遍历所有推文
    for tweet_idx, word_idx in word_positions[word]:
        tweet = tweet_tokens[tweet_idx]
        # 显式判断左右都有上下文,避免异常开销
        if word_idx - 1 >= 0 and word_idx + 1 < len(tweet):
            left_word = tweet[word_idx-1]
            right_word = tweet[word_idx+1]
            # 左右词也必须在词向量表中才能计算
            if left_word in word2vec and right_word in word2vec:
                # 直接相加除以2,比调用np.mean效率更高
                context_vec = (word2vec.get_vector(left_word) + word2vec.get_vector(right_word)) / 2
                context_vecs.append(context_vec)
    # 所有上下文遍历完成后再求均值,不用在循环内判断是否是最后一条推文
    if context_vecs:
        vectors[word] = np.mean(context_vecs, axis=0)
    else:
        # 没有有效上下文的词可以随机初始化或者用零向量,根据需求调整
        vectors[word] = np.random.uniform(-0.25, 0.25, emb_dim)

第三步:可选并行加速

因为每个未知词的向量计算完全独立,可以用多进程进一步提速,CPU核心数充足的情况下速度可以提升数倍:

from multiprocessing import Pool
import numpy as np

def calculate_oov_vec(word):
    context_vecs = []
    for tweet_idx, word_idx in word_positions[word]:
        tweet = tweet_tokens[tweet_idx]
        if word_idx - 1 >= 0 and word_idx + 1 < len(tweet):
            left_word = tweet[word_idx-1]
            right_word = tweet[word_idx+1]
            if left_word in word2vec and right_word in word2vec:
                context_vec = (word2vec.get_vector(left_word) + word2vec.get_vector(right_word)) / 2
                context_vecs.append(context_vec)
    if context_vecs:
        return (word, np.mean(context_vecs, axis=0))
    else:
        return (word, np.random.uniform(-0.25, 0.25, emb_dim))

# 进程数选CPU核心数的1-2倍即可
with Pool(processes=8) as pool:
    results = pool.map(calculate_oov_vec, unique_oov_words)
vectors = dict(results)

优化效果

  • 第一步倒排索引的时间复杂度仅为O(总词数),仅需要遍历所有推文一次
  • 去掉重复未知词后,需要计算的词量通常可以减少90%以上
  • 整体速度至少可以提升两个数量级,原方案需要数天的计算量优化后几小时即可完成

内容的提问来源于stack exchange,提问作者Dmitry Sokolov

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.28 06:36:05