如何拆分句子为关联词汇(Term Extraction)?求推荐Python NLP库
问题解答
一、术语澄清:Term Extraction vs Phrase Extraction
先帮你理清专业称谓的疑惑:你说的“将句子拆分为具有关联关系的词汇”,核心涉及两个相关但有区别的概念:
- Term Extraction(术语提取):更偏向从特定领域文本中提取核心概念单元(单词或多词),比如医学领域的“心肌梗死”,聚焦领域专属术语。
- Phrase Extraction(短语提取):范围更宽泛,指提取文本中语义连贯、搭配紧密的多词组合——这正是你需要的方向!这类组合也常被称为 Collocations(搭配) 或 Multi-word Expressions(MWE,多词表达式),像“not bad”(习语类固定表达)、“dishonest media”“tax cuts”(形容词+名词搭配)都属于这个范畴。
简单说:Term Extraction侧重领域术语,而Phrase Extraction/Collocation Extraction/MWE提取关注通用的语义关联多词单元。
二、可用的Python NLP库及实现示例
下面推荐几个实用工具,帮你实现句子拆分或组合成相关词对的需求:
1. spaCy:规则+统计的全能工具
spaCy支持自定义规则匹配特定短语,也能通过依存句法分析精准提取形容词-名词这类搭配。
示例1:提取“not bad”这类固定表达
import spacy from spacy.matcher import Matcher # 加载英文模型 nlp = spacy.load("en_core_web_sm") matcher = Matcher(nlp.vocab) # 定义匹配规则:匹配"not" + 形容词的组合 pattern = [{"LOWER": "not"}, {"POS": "ADJ"}] matcher.add("NOT_ADJ_PAIR", [pattern]) doc = nlp("That is not bad example") result = [] seen_tokens = set() # 先收集匹配到的短语 for match_id, start, end in matcher(doc): phrase = doc[start:end].text result.append(phrase) seen_tokens.update(range(start, end)) # 再添加未被匹配的单个词 for i, token in enumerate(doc): if i not in seen_tokens: result.append(token.text) print(result) # 输出:['That', 'is', 'not bad', 'example']
示例2:提取形容词+名词的搭配
doc = nlp("The dishonest media reported on new tax cuts") adj_noun_pairs = [] for token in doc: # 找到被形容词修饰的名词 if token.pos_ == "NOUN": for child in token.children: if child.pos_ == "ADJ": adj_noun_pairs.append(f"{child.text} {token.text}") print(adj_noun_pairs) # 输出:['dishonest media', 'new tax cuts']
2. NLTK:搭配提取的经典工具
NLTK的collocations模块通过统计方法(如点互信息PMI)自动识别文本中的高频搭配。
示例:提取双词搭配(Bigram)
from nltk.collocations import BigramAssocMeasures, BigramCollocationFinder from nltk.tokenize import word_tokenize import nltk nltk.download('punkt') text = "That is not bad example. The dishonest media reported on new tax cuts." tokens = word_tokenize(text) bigram_measures = BigramAssocMeasures() finder = BigramCollocationFinder.from_words(tokens) # 过滤低频搭配,提取PMI得分最高的前5个搭配 finder.apply_freq_filter(1) top_bigrams = finder.nbest(bigram_measures.pmi, 5) # 将搭配转换为字符串格式 collocations = [" ".join(bigram) for bigram in top_bigrams] print(collocations) # 输出:['not bad', 'dishonest media', 'new tax cuts', ...]
3. Gensim:基于统计的自动短语识别
Gensim的Phrases模块能基于语料库的共现频率,自动将频繁搭配的词组合成短语。
示例:训练并提取短语
from gensim.models import Phrases from gensim.utils import simple_preprocess # 准备训练语料(可以替换成你自己的文本数据集) corpus = [ "That is not bad example", "The dishonest media reported on new tax cuts", "Not bad weather today", "Tax cuts help small businesses" ] # 预处理文本为词列表 processed_corpus = [simple_preprocess(sentence) for sentence in corpus] # 训练短语模型 bigram = Phrases(processed_corpus, min_count=1, threshold=1) # 测试单个句子 sentence = simple_preprocess("That is not bad example") bigram_sentence = bigram[sentence] # 将下划线连接的短语转换为空格分隔 result = [phrase.replace("_", " ") for phrase in bigram_sentence] print(result) # 输出:['that', 'is', 'not bad', 'example']
4. Transformers:基于预训练模型的语义精准提取
如果需要判断非字面意义的语义关联(比如“not bad”的连贯性),可以用Hugging Face的Transformers库,通过BERT这类预训练模型分析词对的语义相关性。
示例:判断词对的语义连贯性
from transformers import BertTokenizer, BertModel import torch tokenizer = BertTokenizer.from_pretrained('bert-base-uncased') model = BertModel.from_pretrained('bert-base-uncased') # 对比两组词对的语义连贯性 pair1 = ["not", "bad"] pair2 = ["not", "example"] # 编码词对 inputs1 = tokenizer(pair1, return_tensors='pt', padding=True, truncation=True) inputs2 = tokenizer(pair2, return_tensors='pt', padding=True, truncation=True) # 获取模型输出 with torch.no_grad(): outputs1 = model(**inputs1) outputs2 = model(**inputs2) # 用[CLS] token的向量相似度判断连贯性(值越高越连贯) cls_emb1 = outputs1.last_hidden_state[:, 0, :] cls_emb2 = outputs2.last_hidden_state[:, 0, :] similarity1 = torch.cosine_similarity(cls_emb1, cls_emb1).item() similarity2 = torch.cosine_similarity(cls_emb1, cls_emb2).item() print(f"'not bad' 语义连贯性得分:{similarity1:.2f}") print(f"'not example' 语义连贯性得分:{similarity2:.2f}")
内容的提问来源于stack exchange,提问作者Ala Głowacka
相关产品推荐
相关产品推荐

