You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用spaCy提取字符串词元?优化pandas中spaCy词形还原效率

问题1:如何使用spaCy获取字符串中的所有词元(lemma)?
  • 先加载spaCy对应语言的预训练模型,用模型处理目标文本得到Doc对象
  • 遍历Doc中的每个Token,提取其lemma_属性即可
  • 可选:过滤停用词、标点符号等无意义的token

示例代码:

import spacy

# 加载英文预训练模型,中文可替换为zh_core_web_sm
nlp = spacy.load("en_core_web_sm")

text = "I am running in the park and I ran yesterday"
doc = nlp(text)

# 获取所有词元
all_lemmas = [token.lemma_ for token in doc]
print(all_lemmas)  # 输出:['I', 'be', 'run', 'in', 'the', 'park', 'and', 'I', 'run', 'yesterday']

# 过滤停用词和标点后的词元
filtered_lemmas = [token.lemma_ for token in doc if not token.is_stop and not token.is_punct]
print(filtered_lemmas)  # 输出:['run', 'park', 'run', 'yesterday']
问题2:优化pandas DataFrame中文本词形还原的效率

你的原函数运行慢主要有两个原因:

  1. 逐行调用nlp(text)没有利用spaCy的批量处理优化,每次处理都有额外初始化开销
  2. 循环拼接字符串line = line + word.lemma_ + " "效率极低——Python字符串是不可变对象,每次拼接都会生成新的字符串

优化方案:

  1. 使用spaCy的nlp.pipe()方法批量处理文本,这是spaCy专为批量任务设计的接口,支持多进程加速,能大幅提升处理速度
  2. 用str.join()代替循环拼接字符串,一次性生成结果,避免频繁创建新字符串对象

示例代码:

import spacy
import pandas as pd

# 加载模型
nlp = spacy.load("en_core_web_sm")

# 构造示例DataFrame
df = pd.DataFrame({
    "text": [
        "I am running in the park",
        "She runs every morning",
        "They ran last week",
        "We will run tomorrow"
    ]
})

# 定义处理单个Doc的函数
def get_lemmas(doc):
    # 可选过滤停用词和标点,直接拼接所有词元则去掉if判断
    return " ".join([token.lemma_ for token in doc if not token.is_stop and not token.is_punct])

# 批量处理整个文本列:batch_size根据数据量调整,n_process=-1使用所有CPU核心
df["lemmatized_text"] = [get_lemmas(doc) for doc in nlp.pipe(df["text"], batch_size=50, n_process=-1)]

print(df["lemmatized_text"])

关键优化点说明:

  • nlp.pipe()会将文本批量传入模型处理,减少重复初始化开销,比逐行调用nlp(text)效率提升数倍
  • str.join()先收集所有词元到列表,再一次性拼接成字符串,比循环拼接效率高得多,尤其是处理长文本时
  • n_process=-1开启多进程处理,能充分利用CPU资源进一步缩短时间(Windows系统若出现问题可设为1)

内容的提问来源于stack exchange,提问作者Patrick

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.17 06:25:01