You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何通过排除重复词加速Stanza词形还原器处理大文本?

针对大语料优化Stanza词形还原效率的方案

给定条件

我有一份小型示例文档,内容如下:

d ='''
I go to school by the school bus everyday with all of my best friends. 
There are several students who also take the buses to school. Buses are quite cheap in my city.
The city which I live in has an enormous number of brilliant schools with smart students.
We have a nice math teacher in my school whose name is Jane Doe.
She also teaches several other topics in our school, including physics, chemistry and sometimes literature as a substitute teacher.
Other classes don't appreciate her efforts as much as my class. She must be nominated as the best school's teacher.
My school is located far from my apartment. This is why, I am taking the bus to school everyday.
'''

需求目标

针对实际场景中4000~8000词的大文本,希望通过跳过已完成词形还原的重复词来加速Stanza词形还原器,且不能用set()提取唯一词形,必须保留结果中的重复词形。

重复词处理规则

示例中重复词的处理逻辑如下:

Word                 Lemma
--------------------------------------------------
school               school
school               school <<<<<< 重复,直接复用已有还原结果
bus                  bus
everyday             everyday
friends              friend
students             student
buses                bus
school               school
Buses                bus <<<<<< 重复,直接复用已有还原结果
cheap                cheap
city                 city
city                 city <<<<<< 重复,直接复用已有还原结果
live                 live
enormous             enormous
number               number
brilliant            brilliant
schools              school
smart                smart
students             student
nice                 nice
math                 math
teacher              teacher
school               school <<<<<< 重复,直接复用已有还原结果
Jane                 jane
Doe                  doe
teaches              teach
topics               topic
school               school <<<<<< 重复,直接复用已有还原结果
including            include
physics              physics
chemistry            chemistry
literature           literature
substitute           substitute
teacher              teacher <<<<<< 重复,直接复用已有还原结果
classes              class
appreciate           appreciate
efforts              effort
class                class
nominated            nominate
school               school <<<<<< 重复,直接复用已有还原结果
teacher              teacher
school               school <<<<<< 重复,直接复用已有还原结果
located              locate
apartment            apartment
bus                  bus
school               school <<<<<< 重复,直接复用已有还原结果
everyday             everyday <<<<<< 重复,直接复用已有还原结果

现有低效实现方案

import stanza
import nltk
nltk_modules = ['punkt',
                'averaged_perceptron_tagger',
                'stopwords',
                'wordnet',
                'omw-1.4',
               ]
nltk.download(nltk_modules, quiet=True, raise_on_error=True,)
STOPWORDS = nltk.corpus.stopwords.words(nltk.corpus.stopwords.fileids())

nlp = stanza.Pipeline(lang='en', processors='tokenize,lemma,pos', tokenize_no_ssplit=True,download_method=DownloadMethod.REUSE_RESOURCES)
doc = nlp(d)
%timeit -n 10000 [ wlm.lower() for _, s in enumerate(doc.sentences) for _, w in enumerate(s.words) if (wlm:=w.lemma) and len(wlm)>2 and wlm not in STOPWORDS]
# 执行效率:10.5 ms ± 112 µs per loop (mean ± std. dev. of 7 runs, 10000 loops each)

替代方案(仍有优化空间)

def get_lm():
  words_list = list()
  lemmas_list = list()
  for _, vsnt in enumerate(doc.sentences):
    for _, vw in enumerate(vsnt.words):
      wlm = vw.lemma.lower()
      wtxt = vw.text.lower()
      if wtxt in words_list and wlm in lemmas_list:
        lemmas_list.append(wlm)
      elif ( wtxt not in words_list and wlm and len(wlm) > 2 and wlm not in STOPWORDS ):
        lemmas_list.append(wlm)
      words_list.append(wtxt)
  return lemmas_list
%timeit -n 10000 get_lm()
# 执行效率:7.85 ms ± 66.6 µs per loop (mean ± std. dev. of 7 runs, 10000 loops each)

预期输出结果

处理示例文档后需保留重复词形,最终结果如下:

lm = [ wlm.lower() for _, s in enumerate(doc.sentences) for _, w in enumerate(s.words) if (wlm:=w.lemma) and len(wlm)>2 and wlm not in STOPWORDS] # 方案1
# lm = get_lm() # 方案2
print(len(lm), lm)
# 输出:47 ['school', 'school', 'bus', 'everyday', 'friend', 'student', 'bus', 'school', 'bus', 'cheap', 'city', 'city', 'live', 'enormous', 'number', 'brilliant', 'school', 'smart', 'student', 'nice', 'math', 'teacher', 'school', 'jane', 'doe', 'teach', 'topic', 'school', 'include', 'physics', 'chemistry', 'literature', 'substitute', 'teacher', 'class', 'appreciate', 'effort', 'class', 'nominate', 'school', 'teacher', 'school', 'locate', 'apartment', 'bus', 'school', 'everyday']

优化疑问

针对4000~8000词的大语料或文档,有没有更高效的实现方案?


优化方案推荐

核心思路是用字典缓存已完成还原的词形映射,因为字典的in操作是O(1)时间复杂度,远快于列表的O(n)查找,这在大语料下能带来显著性能提升。

实现代码如下:

import stanza
import nltk

# 下载NLTK资源
nltk_modules = ['punkt', 'averaged_perceptron_tagger', 'stopwords', 'wordnet', 'omw-1.4']
nltk.download(nltk_modules, quiet=True, raise_on_error=True)
STOPWORDS = set(nltk.corpus.stopwords.words(nltk.corpus.stopwords.fileids()))  # 转成集合,提升查找速度

# 初始化Stanza管道
nlp = stanza.Pipeline(lang='en', processors='tokenize,lemma,pos', tokenize_no_ssplit=True, download_method=stanza.DownloadMethod.REUSE_RESOURCES)
doc = nlp(d)

def optimized_lemmatize():
    lemma_cache = {}  # 缓存:小写词 -> 对应的lemma
    result = []
    for sent in doc.sentences:
        for word in sent.words:
            word_lower = word.text.lower()
            # 先检查缓存
            if word_lower in lemma_cache:
                lemma = lemma_cache[word_lower]
            else:
                # 首次处理,获取lemma并缓存
                lemma = word.lemma.lower()
                # 仅缓存符合条件的lemma(避免缓存停用词或短词)
                if lemma and len(lemma) > 2 and lemma not in STOPWORDS:
                    lemma_cache[word_lower] = lemma
            # 符合条件则加入结果
            if lemma and len(lemma) > 2 and lemma not in STOPWORDS:
                result.append(lemma)
    return result

# 测试效率
%timeit -n 10000 optimized_lemmatize()

优化点说明

  1. 缓存结构优化:用字典lemma_cache存储已处理的小写词到lemma的映射,查找时间从列表的O(n)降到O(1),大语料下性能提升明显。
  2. 停用词集合化:把STOPWORDS从列表转成集合,in操作同样从O(n)变为O(1)。
  3. 逻辑简化:避免冗余的列表查找和判断,逻辑更清晰,减少不必要的计算。
  4. 缓存过滤:仅缓存符合条件的lemma(非停用词、长度>2),减少缓存体积,提升后续查找效率。

这个方案在大语料下的性能会远优于之前的两种实现,同时完全满足保留重复词形的需求。


内容的提问来源于stack exchange,提问作者farid

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.16 15:17:07