计算Trigram概率时data.table赋值错误及警告问题求助
I've run into this exact length mismatch issue with data.table when implementing Kneser-Kney smoothing before—let's break down what's happening and fix it step by step.
What's Causing the Error?
Your error comes down to a mismatch between the number of values generated by your probability calculation (13867 items) and the number of rows in your tri_words table (3932). When you use n1w12[list(word_1, word_2), .N] or bi_words[list(word_1, word_2), ...], you're not properly aligning the lookup results to each row in tri_words. Data.table is returning results based on matches in the lookup table, not one value per row in tri_words, leading to length discrepancies.
Step-by-Step Fix
Let's restructure your code to use data.table's join functionality (which ensures row alignment) instead of direct indexing. We'll assume you already have:
tri_words: A data.table with columnsword_1,word_2,word_3,count(trigram counts)bi_words: A data.table with columnsword_1,word_2,count(bigram counts)uni_words: A data.table for unigram counts (for backoff if needed)
1. Set Proper Keys for Joins
First, make sure all tables have keys defined to enable fast, aligned joins:
library(data.table) library(quanteda) # Set keys for efficient joining setkey(tri_words, word_1, word_2, word_3) setkey(bi_words, word_1, word_2)
2. Precompute Required Metrics for Each Bigram
Kneser-Kney needs two key values per bigram (w1, w2):
Cn2: The count of the bigram(w1, w2)(frombi_words)n_distinct: The number of unique words that follow(w1, w2)(i.e., uniquew3values for the trigram)
Compute these and join them directly to tri_words:
# Calculate number of unique follow words for each (w1, w2) n1w12 <- tri_words[, .(n_distinct = uniqueN(word_3)), by = .(word_1, word_2)] setkey(n1w12, word_1, word_2) # Join bigram counts and distinct follow words to trigrams tri_words <- tri_words[bi_words, Cn2 := i.count, on = .(word_1, word_2)] tri_words <- tri_words[n1w12, n_distinct := i.n_distinct, on = .(word_1, word_2)]
3. Prepare Backoff Probabilities
Kneser-Kney uses backoff to lower-order n-grams when trigram counts are low. For trigrams, we need the bigram probability P(w3 | w2). Let's compute or join this:
# First, create a reversed bigram table to get P(w3 | w2) (instead of P(w2 | w1)) bi_words_rev <- bi_words[, .(word_1 = word_2, word_2 = word_3, count = count)] setkey(bi_words_rev, word_1, word_2) # Calculate bigram probabilities (using maximum likelihood for simplicity) bi_words_rev[, total := sum(count), by = word_1] bi_words_rev[, P_backoff := count / total] # Join backoff probabilities to trigrams tri_words <- tri_words[bi_words_rev, P_backoff := i.P_backoff, on = .(word_2 = word_1, word_3 = word_2)]
4. Calculate Kneser-Kney Trigram Probabilities
Now we can compute the probability for each trigram row without length mismatches:
discount_value <- 0.75 # Standard Kneser-Kney discount tri_words[, Prob := pmax(count - discount_value, 0) / Cn2 + (discount_value / Cn2) * n_distinct * P_backoff]
Why This Works
By using data.table's on parameter for joins, we ensure every row in tri_words gets exactly one corresponding value for Cn2, n_distinct, and P_backoff. This eliminates the length mismatch because all calculation vectors are the same length as tri_words (3932 rows), so assigning to the Prob column works perfectly.
The original warnings about non-multiple lengths were also caused by misaligned vectors—this fix resolves those too.
内容的提问来源于stack exchange,提问作者ch.elahe

