You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

计算Trigram概率时data.table赋值错误及警告问题求助

Fixing Kneser-Kney Trigram Probability Calculation in data.table

I've run into this exact length mismatch issue with data.table when implementing Kneser-Kney smoothing before—let's break down what's happening and fix it step by step.

What's Causing the Error?

Your error comes down to a mismatch between the number of values generated by your probability calculation (13867 items) and the number of rows in your tri_words table (3932). When you use n1w12[list(word_1, word_2), .N] or bi_words[list(word_1, word_2), ...], you're not properly aligning the lookup results to each row in tri_words. Data.table is returning results based on matches in the lookup table, not one value per row in tri_words, leading to length discrepancies.

Step-by-Step Fix

Let's restructure your code to use data.table's join functionality (which ensures row alignment) instead of direct indexing. We'll assume you already have:

  • tri_words: A data.table with columns word_1, word_2, word_3, count (trigram counts)
  • bi_words: A data.table with columns word_1, word_2, count (bigram counts)
  • uni_words: A data.table for unigram counts (for backoff if needed)

1. Set Proper Keys for Joins

First, make sure all tables have keys defined to enable fast, aligned joins:

library(data.table)
library(quanteda)

# Set keys for efficient joining
setkey(tri_words, word_1, word_2, word_3)
setkey(bi_words, word_1, word_2)

2. Precompute Required Metrics for Each Bigram

Kneser-Kney needs two key values per bigram (w1, w2):

  • Cn2: The count of the bigram (w1, w2) (from bi_words)
  • n_distinct: The number of unique words that follow (w1, w2) (i.e., unique w3 values for the trigram)

Compute these and join them directly to tri_words:

# Calculate number of unique follow words for each (w1, w2)
n1w12 <- tri_words[, .(n_distinct = uniqueN(word_3)), by = .(word_1, word_2)]
setkey(n1w12, word_1, word_2)

# Join bigram counts and distinct follow words to trigrams
tri_words <- tri_words[bi_words, Cn2 := i.count, on = .(word_1, word_2)]
tri_words <- tri_words[n1w12, n_distinct := i.n_distinct, on = .(word_1, word_2)]

3. Prepare Backoff Probabilities

Kneser-Kney uses backoff to lower-order n-grams when trigram counts are low. For trigrams, we need the bigram probability P(w3 | w2). Let's compute or join this:

# First, create a reversed bigram table to get P(w3 | w2) (instead of P(w2 | w1))
bi_words_rev <- bi_words[, .(word_1 = word_2, word_2 = word_3, count = count)]
setkey(bi_words_rev, word_1, word_2)

# Calculate bigram probabilities (using maximum likelihood for simplicity)
bi_words_rev[, total := sum(count), by = word_1]
bi_words_rev[, P_backoff := count / total]

# Join backoff probabilities to trigrams
tri_words <- tri_words[bi_words_rev, P_backoff := i.P_backoff, on = .(word_2 = word_1, word_3 = word_2)]

4. Calculate Kneser-Kney Trigram Probabilities

Now we can compute the probability for each trigram row without length mismatches:

discount_value <- 0.75  # Standard Kneser-Kney discount

tri_words[, Prob := pmax(count - discount_value, 0) / Cn2 + 
            (discount_value / Cn2) * n_distinct * P_backoff]

Why This Works

By using data.table's on parameter for joins, we ensure every row in tri_words gets exactly one corresponding value for Cn2, n_distinct, and P_backoff. This eliminates the length mismatch because all calculation vectors are the same length as tri_words (3932 rows), so assigning to the Prob column works perfectly.

The original warnings about non-multiple lengths were also caused by misaligned vectors—this fix resolves those too.

内容的提问来源于stack exchange,提问作者ch.elahe

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 23:02:47