You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何加速Python处理CSV评论文本去停用词的运行速度?

优化停用词移除的运行时间

我正尝试从.csv文件的'reviews.text'列中移除停用词,运行代码后耗时10分钟,如何缩短运行时间?

原代码

import pandas as pd
from os import chdir, path
import spacy
from spacytextblob.spacytextblob import SpacyTextBlob

nlp = spacy.load('en_core_web_sm')
nlp.add_pipe('spacytextblob')

chdir(path.dirname(__file__))

file_path = 'amazon_product_reviews.csv'
dataframe = pd.read_csv(file_path, dtype={'id': str, 'name': str, 'asins': str, 'brand': str, 'categories': str, 'keys': str, 'manufacturer': str, 'reviews.date': str, 'reviews.dateAdded': str, 'reviews.dateSeen': str, 'reviews.didPurchase': str, 'reviews.doRecommend': str, 'reviews.id': str, 'reviews.numHelpful': str, 'reviews.rating': str, 'reviews.sourceURLs': str, 'reviews.text': str, 'reviews.title': str, 'reviews.userCity': str, 'reviews.userProvince': str, 'reviews.username': str, })

reviews_data = dataframe['reviews.text']

clean_data = dataframe.dropna(subset=['reviews.text'])

def preprocess_text(text):
    doc = nlp(text)
    
    cleaned_tokens = [token.text.lower() for token in doc if token.is_alpha and not token.is_stop]
    
    cleaned_text = ' '.join(cleaned_tokens)
    
    return cleaned_text

clean_data = clean_data.copy()
clean_data['processed_reviews'] = clean_data['reviews.text'].apply(preprocess_text)

print("Cleaned Data:")
print(clean_data[['reviews.text', 'processed_reviews']].head())

cProfile性能分析结果

我运行了cProfile来查看代码中耗时最长的部分,结果如下:

302681427 function calls (296741014 primitive calls) in 294.594 seconds

   Ordered by: cumulative time

   ncalls  tottime  percall  cumtime  percall filename:lineno(function)
     10/1    0.000    0.000  294.659  294.659 {built-in method builtins.exec}
        1    0.003    0.003  294.639  294.639 test3.py:10(main)
        1    0.000    0.000  293.915  293.915 series.py:4769(apply)
        1    0.000    0.000  293.915  293.915 apply.py:1409(apply)
        1    0.000    0.000  293.915  293.915 apply.py:1482(apply_standard)
        1    0.000    0.000  293.915  293.915 base.py:891(_map_values)
        1    0.121    0.121  293.915  293.915 algorithms.py:1667(map_array)
    34659    0.047    0.000  293.793    0.008 test3.py:29(preprocess_text)
    34659    0.465    0.000  293.253    0.008 language.py:1016(__call__)
   138636   32.197    0.000  242.236    0.002 trainable_pipe.pyx:40(__call__)
   138636    0.531    0.000  205.376    0.001 model.py:330(predict)
4678965/277272    1.998    0.000  203.319    0.001 model.py:307(__call__)
1628973/138636    2.245    0.000  187.239    0.001 chain.py:48(forward)
   242613    0.263    0.000  180.916    0.001 with_array.py:32(forward)
   519885  157.488    0.000  157.731    0.000 numpy_ops.pyx:91(gemm)
   346590    2.671    0.000  145.591    0.000 maxout.py:45(forward)
   103977    0.291    0.000  132.341    0.001 with_array.py:70(_list_forward)
   277272    0.548    0.000  127.896    0.000 residual.py:28(forward)
    69318    0.632    0.000  107.110    0.002 tb_framework.py:33(forward)

优化方案

  • 精简Spacy组件:你加载了完整的en_core_web_sm模型,但只用到分词、词性判断和停用词过滤。加载时禁用不需要的组件,同时删掉没用的spacytextblob:

    nlp = spacy.load('en_core_web_sm', disable=['parser', 'ner', 'textcat'])
    # 删掉这行:nlp.add_pipe('spacytextblob')
    

    这能大幅减少模型的计算开销。

  • 用批量处理替代逐行apply:Spacy的nlp.pipe()支持批量处理文本,还能多进程加速。把apply替换成:

    clean_data['processed_reviews'] = [
        ' '.join([token.text.lower() for token in doc if token.is_alpha and not token.is_stop])
        for doc in nlp.pipe(clean_data['reviews.text'], batch_size=1000, n_process=-1)
    ]
    

    batch_size可根据内存调整,n_process=-1会调用所有CPU核心,效率比逐行处理高很多。

  • 减少不必要的数据操作:删除没用的reviews_data = dataframe['reviews.text']行,去掉clean_data = clean_data.copy(),避免多余的内存复制。

  • 换用轻量工具:如果只需要停用词过滤,没必要用Spacy的重型模型。用NLTK实现更简单快速:

    import nltk
    from nltk.corpus import stopwords
    from nltk.tokenize import word_tokenize
    
    # 第一次运行需要下载资源
    nltk.download('stopwords')
    nltk.download('punkt')
    stop_words = set(stopwords.words('english'))
    
    def preprocess_text(text):
        tokens = word_tokenize(text.lower())
        return ' '.join([token for token in tokens if token.isalpha() and token not in stop_words])
    
  • 优化CSV读取:你给所有列指定了str类型,还加载了无关列。只读取需要的reviews.text列,减少内存占用:

    dataframe = pd.read_csv(file_path, usecols=['reviews.text'], dtype={'reviews.text': str})
    

内容的提问来源于stack exchange,提问作者Huy Dang

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.29 08:14:54