You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于带权重unigram列表生成有意义的bigram?

解决方案

要生成所有可能的bigram并筛选合理组合,推荐以下几种实用方法:

方法1:基于词性规则的语法过滤

先给每个单词标注词性,再根据语法搭配规则(比如形容词+动名词、动词+名词这类常见合理组合)筛选。这种方法简单高效,不需要额外语料或大模型。

步骤:

  1. 用NLTK给所有unigram做词性标注
  2. 定义你认为合理的词性组合(比如形容词+动名词、动名词+名词等,可根据需求自定义)
  3. 生成所有两两组合的bigram,用词性规则过滤掉不符合的

代码示例:

import nltk
from itertools import product

# 你的带权重unigram列表
unigrams = [('bottom', 507.95), ('straight', 426.5), ('comment', 415.5), ('wearing', 398.55), ('room', 397.85), ('wondering', 396.85), ('difficult', 382.85), ('sleeping', 381.65), ('comments', 381.1), ('looked', 379.0), ('interest', 378.2), ('missing', 373.5), ('harder', 373.1), ('planning', 370.05), ('answer', 367.15), ('allowed', 364.85), ('bunch', 361.0), ('recommend', 360.45), ('worst', 359.3), ('technically', 359.15)]

# 提取单词列表
words = [word for word, _ in unigrams]

# 词性标注(首次运行需下载nltk的punkt和averaged_perceptron_tagger)
# nltk.download('punkt')
# nltk.download('averaged_perceptron_tagger')
pos_tags = nltk.pos_tag(words)
word_to_pos = {word: tag for word, tag in pos_tags}

# 自定义合法词性组合(可根据需求增减)
valid_pos_pairs = [
    # 形容词+动名词(比如difficult sleeping)
    ('JJ', 'VBG'), ('JJR', 'VBG'), ('JJS', 'VBG'),
    # 动名词+名词(比如wearing room)
    ('VBG', 'NN'), ('VBG', 'NNS'),
    # 动词过去式+名词(比如looked comment)
    ('VBD', 'NN'), ('VBD', 'NNS'),
    # 副词+形容词(比如technically difficult)
    ('RB', 'JJ'), ('RB', 'JJR'),
    # 名词+名词(比如comment comments)
    ('NN', 'NN'), ('NN', 'NNS')
]

# 生成所有可能的bigram并筛选
all_possible_bigrams = product(words, repeat=2)
valid_bigrams = []
for w1, w2 in all_possible_bigrams:
    if (word_to_pos.get(w1), word_to_pos.get(w2)) in valid_pos_pairs:
        valid_bigrams.append((w1, w2))

# 输出部分结果
print("筛选后的合理bigram示例:", valid_bigrams[:10])

方法2:基于预训练语言模型的语义判断

用预训练的语言模型(比如BERT)计算每个bigram的困惑度(perplexity),困惑度越低说明这个组合在语义上越通顺自然。这种方法能更好地处理语法规则覆盖不到的语义合理性问题。

代码示例:

from transformers import BertTokenizer, BertForMaskedLM
import torch
from itertools import product

unigrams = [('bottom', 507.95), ('straight', 426.5), ('comment', 415.5), ('wearing', 398.55), ('room', 397.85), ('wondering', 396.85), ('difficult', 382.85), ('sleeping', 381.65), ('comments', 381.1), ('looked', 379.0), ('interest', 378.2), ('missing', 373.5), ('harder', 373.1), ('planning', 370.05), ('answer', 367.15), ('allowed', 364.85), ('bunch', 361.0), ('recommend', 360.45), ('worst', 359.3), ('technically', 359.15)]
words = [word for word, _ in unigrams]

# 加载预训练模型和分词器
tokenizer = BertTokenizer.from_pretrained('bert-base-uncased')
model = BertForMaskedLM.from_pretrained('bert-base-uncased')
model.eval()

# 计算单个bigram的困惑度
def get_perplexity(bigram):
    text = ' '.join(bigram)
    inputs = tokenizer(text, return_tensors='pt')
    with torch.no_grad():
        outputs = model(**inputs, labels=inputs['input_ids'])
        loss = outputs.loss
    return torch.exp(loss).item()

# 生成所有bigram并按困惑度筛选
all_bigrams = product(words, repeat=2)
# 设定困惑度阈值(值越小越合理,可根据实际情况调整)
threshold = 80
valid_bigrams = []
for bg in all_bigrams:
    ppl = get_perplexity(bg)
    if ppl < threshold:
        valid_bigrams.append((bg, ppl))

# 按困惑度从小到大排序(越靠前越合理)
valid_bigrams.sort(key=lambda x: x[1])
print("语义合理的bigram示例:", [bg for bg, _ in valid_bigrams[:10]])

方法3:基于互信息的统计筛选(需语料支持)

如果你有相关领域的语料库,可以计算两个词的互信息(Mutual Information),互信息越高说明这两个词在语料中越常搭配出现,组合越合理。

代码示例(假设已有语料):

from itertools import product
from collections import defaultdict
import math

# 替换为你的实际语料库
corpus = [
    "difficult sleeping", "wearing room", "looked comment",
    "missing interest", "planning answer", "allowed bunch",
    "worst comment", "technically difficult", "straight answer"
    # 更多语料内容...
]

# 统计词频和bigram出现次数
word_counts = defaultdict(int)
bigram_counts = defaultdict(int)
for sentence in corpus:
    tokens = sentence.split()
    for token in tokens:
        word_counts[token] += 1
    for i in range(len(tokens)-1):
        bg = (tokens[i], tokens[i+1])
        bigram_counts[bg] += 1

total_words = sum(word_counts.values())
total_bigrams = sum(bigram_counts.values())

# 计算互信息
def calculate_mi(w1, w2):
    p_w1 = word_counts.get(w1, 0) / total_words if total_words else 0
    p_w2 = word_counts.get(w2, 0) / total_words if total_words else 0
    p_bg = bigram_counts.get((w1, w2), 0) / total_bigrams if total_bigrams else 0
    if p_w1 == 0 or p_w2 == 0 or p_bg == 0:
        return 0
    return math.log2(p_bg / (p_w1 * p_w2))

# 你的unigram列表
unigrams = [('bottom', 507.95), ('straight', 426.5), ('comment', 415.5), ('wearing', 398.55), ('room', 397.85), ('wondering', 396.85), ('difficult', 382.85), ('sleeping', 381.65), ('comments', 381.1), ('looked', 379.0), ('interest', 378.2), ('missing', 373.5), ('harder', 373.1), ('planning', 370.05), ('answer', 367.15), ('allowed', 364.85), ('bunch', 361.0), ('recommend', 360.45), ('worst', 359.3), ('technically', 359.15)]
words = [word for word, _ in unigrams]

# 生成所有bigram并筛选
all_bigrams = product(words, repeat=2)
# 设定互信息阈值(值越大越合理)
threshold = 0.3
valid_bigrams = []
for bg in all_bigrams:
    mi = calculate_mi(*bg)
    if mi > threshold:
        valid_bigrams.append((bg, mi))

# 按互信息从大到小排序
valid_bigrams.sort(key=lambda x: x[1], reverse=True)
print("统计上合理的bigram示例:", [bg for bg, _ in valid_bigrams[:10]])

内容的提问来源于stack exchange,提问作者Mara de Jess Garcia Santiago

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.11 05:40:22