如何使用Python从句子中提取名词性复合词?
提取名词性复合词的解决方案
下面提供几种实用方法,帮你用NLTK提取“basketball shoe”这类名词组合:
方法1:自定义正则分块器
通过定义规则匹配连续的名词词性标签,将它们合并为名词短语:import nltk from nltk.chunk import RegexpParser sentence = "I bought a basketball shoe and a coffee table yesterday" words = nltk.word_tokenize(sentence) tags = nltk.pos_tag(words) # 规则:匹配任意数量连续的名词标签(NN、NNP、NNS等) chunk_rule = r""" NP: {<NN.*>+} """ chunk_parser = RegexpParser(chunk_rule) parsed_result = chunk_parser.parse(tags) # 提取所有名词短语 noun_phrases = [] for subtree in parsed_result.subtrees(filter=lambda t: t.label() == 'NP'): phrase = ' '.join(word for word, tag in subtree.leaves()) noun_phrases.append(phrase) print(noun_phrases) # 输出: ['basketball shoe', 'coffee table']方法2:使用预训练分块模型
NLTK自带基于树库训练的分块器,能处理更复杂的名词短语结构:import nltk from nltk.corpus import treebank_chunk # 下载并加载预训练分块器 nltk.download('treebank_chunk') chunker = treebank_chunk.chunkers()[0] sentence = "My favorite sports gear is basketball shoe and tennis racket" words = nltk.word_tokenize(sentence) tags = nltk.pos_tag(words) parsed_result = chunker.parse(tags) noun_phrases = [] for subtree in parsed_result.subtrees(filter=lambda t: t.label() == 'NP'): phrase = ' '.join(word for word, tag in subtree.leaves()) noun_phrases.append(phrase) print(noun_phrases)方法3:滑动窗口合并连续名词
逻辑简单直接,遍历词性序列,把连续的名词标签对应的词合并:import nltk sentence = "She placed the wooden dining table near the window" words = nltk.word_tokenize(sentence) tags = nltk.pos_tag(words) noun_phrases = [] current_phrase = [] for word, tag in tags: if tag.startswith('NN'): current_phrase.append(word) else: if current_phrase: noun_phrases.append(' '.join(current_phrase)) current_phrase = [] # 处理最后一组未闭合的名词 if current_phrase: noun_phrases.append(' '.join(current_phrase)) print(noun_phrases) # 输出: ['wooden dining table', 'window']
内容的提问来源于stack exchange,提问作者Alan K
相关产品推荐
相关产品推荐

