如何用spaCy提取字符串词元?优化pandas中spaCy词形还原效率
问题1:如何使用spaCy获取字符串中的所有词元(lemma)?
- 先加载spaCy对应语言的预训练模型,用模型处理目标文本得到
Doc对象 - 遍历
Doc中的每个Token,提取其lemma_属性即可 - 可选:过滤停用词、标点符号等无意义的token
示例代码:
import spacy # 加载英文预训练模型,中文可替换为zh_core_web_sm nlp = spacy.load("en_core_web_sm") text = "I am running in the park and I ran yesterday" doc = nlp(text) # 获取所有词元 all_lemmas = [token.lemma_ for token in doc] print(all_lemmas) # 输出:['I', 'be', 'run', 'in', 'the', 'park', 'and', 'I', 'run', 'yesterday'] # 过滤停用词和标点后的词元 filtered_lemmas = [token.lemma_ for token in doc if not token.is_stop and not token.is_punct] print(filtered_lemmas) # 输出:['run', 'park', 'run', 'yesterday']
问题2:优化pandas DataFrame中文本词形还原的效率
你的原函数运行慢主要有两个原因:
- 逐行调用
nlp(text)没有利用spaCy的批量处理优化,每次处理都有额外初始化开销 - 循环拼接字符串
line = line + word.lemma_ + " "效率极低——Python字符串是不可变对象,每次拼接都会生成新的字符串
优化方案:
- 使用spaCy的
nlp.pipe()方法批量处理文本,这是spaCy专为批量任务设计的接口,支持多进程加速,能大幅提升处理速度 - 用
str.join()代替循环拼接字符串,一次性生成结果,避免频繁创建新字符串对象
示例代码:
import spacy import pandas as pd # 加载模型 nlp = spacy.load("en_core_web_sm") # 构造示例DataFrame df = pd.DataFrame({ "text": [ "I am running in the park", "She runs every morning", "They ran last week", "We will run tomorrow" ] }) # 定义处理单个Doc的函数 def get_lemmas(doc): # 可选过滤停用词和标点,直接拼接所有词元则去掉if判断 return " ".join([token.lemma_ for token in doc if not token.is_stop and not token.is_punct]) # 批量处理整个文本列:batch_size根据数据量调整,n_process=-1使用所有CPU核心 df["lemmatized_text"] = [get_lemmas(doc) for doc in nlp.pipe(df["text"], batch_size=50, n_process=-1)] print(df["lemmatized_text"])
关键优化点说明:
nlp.pipe()会将文本批量传入模型处理,减少重复初始化开销,比逐行调用nlp(text)效率提升数倍str.join()先收集所有词元到列表,再一次性拼接成字符串,比循环拼接效率高得多,尤其是处理长文本时n_process=-1开启多进程处理,能充分利用CPU资源进一步缩短时间(Windows系统若出现问题可设为1)
内容的提问来源于stack exchange,提问作者Patrick
相关产品推荐
相关产品推荐

