如何在Python NLTK中针对特定单词(如cat)进行句法分块?
针对特定单词的NLTK分块实现方案
好问题!直接用<cat>这种写法在NLTK的正则表达式分块器里是行不通的——它的规则语法是**基于词性标签(POS tags)**匹配的,不会把尖括号里的内容当成单词本身识别。不过咱们有两种简单的方式实现针对特定单词的分块需求,下面一步步拆解:
先回顾你的原示例
原代码和执行结果如下:
tagged = [('The', 'DT'), ('cat', 'NN'), ('sat', 'VBD'), ('on', 'IN'), ('the','DT'), ('mat','NN'), ('the', 'DT'), ('dog', 'NN'), ('chewed', 'VBD')] gram = r"""chk: {<DT>?<NN><VBD>}"""
执行后得到分块结果:
the cat sat、the dog chewed
为什么直接写<cat>不生效?
NLTK的RegexpParser规则里,尖括号<>包裹的内容只能是词性标签,它不会去匹配单词文本。所以<cat>会被当成一个不存在的词性标签,自然匹配不到任何内容。
方案1:自定义标签匹配特定单词
我们可以先把目标单词的词性标签替换成一个自定义标签,再基于这个标签写规则:
from nltk.chunk import RegexpParser # 原标注语料 tagged = [('The', 'DT'), ('cat', 'NN'), ('sat', 'VBD'), ('on', 'IN'), ('the','DT'), ('mat','NN'), ('the', 'DT'), ('dog', 'NN'), ('chewed', 'VBD')] # 将目标单词"cat"的标签替换为自定义的<CAT_TAG>(支持大小写不敏感匹配) custom_tagged = [(word, 'CAT_TAG' if word.lower() == 'cat' else tag) for word, tag in tagged] # 基于自定义标签编写分块规则 gram = r"""chk: {<DT>?<CAT_TAG><VBD>}""" chunk_parser = RegexpParser(gram) result = chunk_parser.parse(custom_tagged) # 提取并打印目标分块 for subtree in result.subtrees(filter=lambda t: t.label() == 'chk'): print(' '.join([word for word, tag in subtree.leaves()]))
执行后会精准输出:
The cat sat,只会匹配包含"cat"的那一组。
方案2:先分块再筛选目标单词
如果不想修改原标注语料,也可以先运行通用分块规则,再从结果里筛选包含特定单词的块:
from nltk.chunk import RegexpParser tagged = [('The', 'DT'), ('cat', 'NN'), ('sat', 'VBD'), ('on', 'IN'), ('the','DT'), ('mat','NN'), ('the', 'DT'), ('dog', 'NN'), ('chewed', 'VBD')] gram = r"""chk: {<DT>?<NN><VBD>}""" chunk_parser = RegexpParser(gram) result = chunk_parser.parse(tagged) # 筛选包含"cat"的分块 target_chunks = [] for subtree in result.subtrees(filter=lambda t: t.label() == 'chk'): chunk_words = [word for word, tag in subtree.leaves()] if 'cat' in chunk_words: target_chunks.append(' '.join(chunk_words)) print(target_chunks) # 输出: ['The cat sat']
两种方案都能实现你的需求,你可以根据自己的场景选择更合适的一种~
内容的提问来源于stack exchange,提问作者Mueladavc
相关产品推荐
相关产品推荐

