You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何拆分句子为关联词汇(Term Extraction)?求推荐Python NLP库

问题解答

一、术语澄清:Term Extraction vs Phrase Extraction

先帮你理清专业称谓的疑惑:你说的“将句子拆分为具有关联关系的词汇”,核心涉及两个相关但有区别的概念:

  • Term Extraction(术语提取):更偏向从特定领域文本中提取核心概念单元(单词或多词),比如医学领域的“心肌梗死”,聚焦领域专属术语。
  • Phrase Extraction(短语提取):范围更宽泛,指提取文本中语义连贯、搭配紧密的多词组合——这正是你需要的方向!这类组合也常被称为 Collocations(搭配) 或 Multi-word Expressions(MWE,多词表达式),像“not bad”(习语类固定表达)、“dishonest media”“tax cuts”(形容词+名词搭配)都属于这个范畴。

简单说:Term Extraction侧重领域术语,而Phrase Extraction/Collocation Extraction/MWE提取关注通用的语义关联多词单元。

二、可用的Python NLP库及实现示例

下面推荐几个实用工具,帮你实现句子拆分或组合成相关词对的需求:

1. spaCy:规则+统计的全能工具

spaCy支持自定义规则匹配特定短语,也能通过依存句法分析精准提取形容词-名词这类搭配。

示例1:提取“not bad”这类固定表达

import spacy
from spacy.matcher import Matcher

# 加载英文模型
nlp = spacy.load("en_core_web_sm")
matcher = Matcher(nlp.vocab)

# 定义匹配规则:匹配"not" + 形容词的组合
pattern = [{"LOWER": "not"}, {"POS": "ADJ"}]
matcher.add("NOT_ADJ_PAIR", [pattern])

doc = nlp("That is not bad example")

result = []
seen_tokens = set()

# 先收集匹配到的短语
for match_id, start, end in matcher(doc):
    phrase = doc[start:end].text
    result.append(phrase)
    seen_tokens.update(range(start, end))

# 再添加未被匹配的单个词
for i, token in enumerate(doc):
    if i not in seen_tokens:
        result.append(token.text)

print(result)
# 输出:['That', 'is', 'not bad', 'example']

示例2:提取形容词+名词的搭配

doc = nlp("The dishonest media reported on new tax cuts")

adj_noun_pairs = []
for token in doc:
    # 找到被形容词修饰的名词
    if token.pos_ == "NOUN":
        for child in token.children:
            if child.pos_ == "ADJ":
                adj_noun_pairs.append(f"{child.text} {token.text}")

print(adj_noun_pairs)
# 输出:['dishonest media', 'new tax cuts']

2. NLTK:搭配提取的经典工具

NLTK的collocations模块通过统计方法(如点互信息PMI)自动识别文本中的高频搭配。

示例:提取双词搭配(Bigram)

from nltk.collocations import BigramAssocMeasures, BigramCollocationFinder
from nltk.tokenize import word_tokenize
import nltk

nltk.download('punkt')

text = "That is not bad example. The dishonest media reported on new tax cuts."
tokens = word_tokenize(text)

bigram_measures = BigramAssocMeasures()
finder = BigramCollocationFinder.from_words(tokens)

# 过滤低频搭配,提取PMI得分最高的前5个搭配
finder.apply_freq_filter(1)
top_bigrams = finder.nbest(bigram_measures.pmi, 5)

# 将搭配转换为字符串格式
collocations = [" ".join(bigram) for bigram in top_bigrams]
print(collocations)
# 输出:['not bad', 'dishonest media', 'new tax cuts', ...]

3. Gensim:基于统计的自动短语识别

Gensim的Phrases模块能基于语料库的共现频率,自动将频繁搭配的词组合成短语。

示例:训练并提取短语

from gensim.models import Phrases
from gensim.utils import simple_preprocess

# 准备训练语料(可以替换成你自己的文本数据集)
corpus = [
    "That is not bad example",
    "The dishonest media reported on new tax cuts",
    "Not bad weather today",
    "Tax cuts help small businesses"
]

# 预处理文本为词列表
processed_corpus = [simple_preprocess(sentence) for sentence in corpus]

# 训练短语模型
bigram = Phrases(processed_corpus, min_count=1, threshold=1)

# 测试单个句子
sentence = simple_preprocess("That is not bad example")
bigram_sentence = bigram[sentence]

# 将下划线连接的短语转换为空格分隔
result = [phrase.replace("_", " ") for phrase in bigram_sentence]
print(result)
# 输出:['that', 'is', 'not bad', 'example']

4. Transformers:基于预训练模型的语义精准提取

如果需要判断非字面意义的语义关联(比如“not bad”的连贯性),可以用Hugging Face的Transformers库,通过BERT这类预训练模型分析词对的语义相关性。

示例:判断词对的语义连贯性

from transformers import BertTokenizer, BertModel
import torch

tokenizer = BertTokenizer.from_pretrained('bert-base-uncased')
model = BertModel.from_pretrained('bert-base-uncased')

# 对比两组词对的语义连贯性
pair1 = ["not", "bad"]
pair2 = ["not", "example"]

# 编码词对
inputs1 = tokenizer(pair1, return_tensors='pt', padding=True, truncation=True)
inputs2 = tokenizer(pair2, return_tensors='pt', padding=True, truncation=True)

# 获取模型输出
with torch.no_grad():
    outputs1 = model(**inputs1)
    outputs2 = model(**inputs2)

# 用[CLS] token的向量相似度判断连贯性(值越高越连贯)
cls_emb1 = outputs1.last_hidden_state[:, 0, :]
cls_emb2 = outputs2.last_hidden_state[:, 0, :]

similarity1 = torch.cosine_similarity(cls_emb1, cls_emb1).item()
similarity2 = torch.cosine_similarity(cls_emb1, cls_emb2).item()

print(f"'not bad' 语义连贯性得分:{similarity1:.2f}")
print(f"'not example' 语义连贯性得分:{similarity2:.2f}")

内容的提问来源于stack exchange,提问作者Ala Głowacka

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 09:30:44