如何生成带int结构与维度信息的lda R包格式列表数据?
lda R Package Great question! I’ve worked extensively with Jonathan Chang’s lda package, so let me walk you through exactly how to replicate the structure of built-in datasets like cora.documents.
First, let’s clarify what makes those built-in datasets special: they aren’t just regular lists. They’re corpus objects with two key characteristics:
- Each element in the list is a 2-column integer matrix (this is what gives you the
[,]dimension info you noticed) - The entire list has a custom class attribute:
c("lda.corpus", "list")
Here’s a step-by-step guide to building your own:
Step 1: Understand the Matrix Format for Each Document
Every document in the corpus must be represented as a 2-column integer matrix:
- Column 1: Integer IDs for words (these map to positions in your vocabulary list, starting at 1—no zeros allowed!)
- Column 2: Integer counts of how many times each word appears in the document
This sparse format is efficient for text data, where most words don’t appear in most documents.
Step 2: Build a Minimal Example Manually
Let’s create a tiny corpus from scratch to see how it works:
# First, define your vocabulary (each word maps to a unique ID starting at 1) vocab <- c("machine", "learning", "lda", "topic", "model") # Create matrix for Document 1: contains "machine" (2x), "learning" (1x), "lda" (3x) doc1 <- matrix( c(1, 2, 3, # Word IDs 2, 1, 3), # Word counts ncol = 2, byrow = FALSE ) # Create matrix for Document 2: contains "lda" (1x), "topic" (2x), "model" (1x) doc2 <- matrix( c(3, 4, 5, 1, 2, 1), ncol = 2, byrow = FALSE ) # Combine documents into a list, then set the required class my_corpus <- list(doc1, doc2) class(my_corpus) <- c("lda.corpus", "list") # Check the structure (matches cora.documents!) str(my_corpus)
When you run str(my_corpus), you’ll see each element is labeled as an int [,1:2] matrix—exactly like the built-in data.
Step 3: Convert Real Text Data to This Format
If you’re working with actual text, you’ll need to first process your text into word counts, then convert to the matrix format. Here’s how to do it with the quanteda package (a popular text processing tool):
library(quanteda) # Sample text data texts <- c( "Machine learning is fun, especially lda topic modeling", "Lda is a great method for topic modeling in text analysis" ) # Create a document-feature matrix (DFM) of word counts dfm_obj <- dfm(texts, tolower = TRUE) # Convert DFM to lda-compatible corpus my_corpus <- lapply(1:nrow(dfm_obj), function(doc_idx) { # Get words that appear in the document (non-zero counts) word_ids <- which(dfm_obj[doc_idx, ] > 0) # Get their counts word_counts <- as.integer(dfm_obj[doc_idx, word_ids]) # Combine into a 2-column matrix matrix(c(word_ids, word_counts), ncol = 2) }) # Set the required class attribute class(my_corpus) <- c("lda.corpus", "list")
Key Notes to Avoid Issues
- Word IDs start at 1: The
ldapackage uses 1-based indexing for vocabulary, so never use 0 as a word ID. - Integer types: Ensure both columns of your matrices are integers (not numeric). The
as.integer()call in the example above handles this. - Class attribute is mandatory: Without setting
class(my_corpus) <- c("lda.corpus", "list"), theldapackage functions won’t recognize your data as a valid corpus.
内容的提问来源于stack exchange,提问作者Chris T.

