如何优化Pandas中处理长文本的Spacy名词过滤函数?
针对Pandas+spaCy批量处理文本的优化方案
以下是几种能显著提升6400行文本处理速度的优化方法:
批量调用spaCy管道,避免逐行开销
spaCy的nlp.pipe()原生支持批量文本处理,相比逐行调用nlp()能大幅降低重复初始化的开销,还可配合多进程并行处理。设置合适的batch_size(根据内存调整,比如100-200)和n_process参数即可:import spacy import pandas as pd # 加载模型时禁用无关组件 nlp = spacy.load("en_core_web_sm", disable=["parser", "ner", "lemmatizer"]) def filter_nouns_batch(texts): # 生成符号翻译表,一次性清理所有文本 symbol_trans = str.maketrans('', '', '{}()[].,:;+-*/&|<>=~$1234567890#_%') cleaned_texts = [text.translate(symbol_trans).strip() for text in texts] filtered_nouns = [] # 批量处理,n_process=-1启用所有CPU核心 for doc in nlp.pipe(cleaned_texts, batch_size=100, n_process=-1): nouns = [token.text for token in doc if token.pos_ == "NOUN"] filtered_nouns.append(nouns) return filtered_nouns # 批量处理整个NOTE列 df['filtered_nouns'] = filter_nouns_batch(df['NOTE'].tolist())禁用spaCy无关管道组件
原代码加载了完整的en_core_web_sm模型,但我们仅需**词性标注(POS Tagging)**功能,依存解析、命名实体识别等组件完全多余。禁用这些组件后,模型的内存占用和计算耗时会大幅降低。优化文本预处理逻辑
原代码先拆分单词再逐词清理符号,效率低下。改用str.maketrans生成统一翻译表,一次性处理整段文本的符号,或用正则批量替换,速度更快:import re # 用正则批量替换目标符号 cleaned_text = re.sub(r'[{}()\[\].,:;+\-*/&|<>=~$0-9#_%]', '', text).strip()可选:改用更轻量的NLP库
如果仅需提取名词,可尝试nltk这类轻量库,它的POS标注器初始化和处理速度更快:import nltk from nltk.tokenize import word_tokenize from nltk.tag import pos_tag # 首次运行需下载资源:nltk.download('punkt'), nltk.download('averaged_perceptron_tagger') def filter_nouns_nltk(text): cleaned_text = re.sub(r'[{}()\[\].,:;+\-*/&|<>=~$0-9#_%]', '', text).strip() tokens = word_tokenize(cleaned_text) tagged_tokens = pos_tag(tokens) # NLTK中名词标签以NN开头(NN、NNS、NNP、NNPS) nouns = [word for word, tag in tagged_tokens if tag.startswith('NN')] return nouns df['filtered_nouns'] = df['NOTE'].apply(filter_nouns_nltk)
内容的提问来源于stack exchange,提问作者Rikky Bhai
相关产品推荐
相关产品推荐

