从LDA预测结果的嵌套列表中获取最频繁主题值
First, since R doesn’t have a built-in function to calculate the mode (most frequent value) of a vector, we’ll create a simple helper function to handle this. It’ll return the most frequent integer in a given vector, and in case of ties, it picks the first value that hits the highest frequency.
# Helper function to compute the mode (most frequent value) get_mode <- function(x) { # Get unique values in the vector unique_vals <- unique(x) # Count occurrences of each unique value freq_counts <- tabulate(match(x, unique_vals)) # Return the unique value with the highest count unique_vals[which.max(freq_counts)] }
Next, apply this function to every sublist in your assignments list. Use sapply() if you want a clean vector of results (ideal for document-level topics) or lapply() if you prefer a list output:
# Your example assignments list assignments <- list( as.integer(c(1, 1, 1, 1, 1, 1, 2, 2, 2, 3, 3)), as.integer(c(1, 1, 1, 2, 2, 2, 2, 2, 3, 3)), as.integer(c(1, 3, 3, 3, 3, 3, 3, 2, 2)) ) # Get dominant topic for each document document_dominant_topics <- sapply(assignments, get_mode) # View the result document_dominant_topics # Output: [1] 1 2 3
Handling Ties (Optional)
If you have a document where two or more topics are equally frequent, the above function returns the first one encountered. If you want to capture all tied modes instead, modify the helper function like this:
get_all_modes <- function(x) { unique_vals <- unique(x) freq_counts <- tabulate(match(x, unique_vals)) max_freq <- max(freq_counts) unique_vals[freq_counts == max_freq] } # Example with a tied document assignments_with_tie <- list(as.integer(c(1,1,2,2))) sapply(assignments_with_tie, get_all_modes) # Output: [1] 1 2
This will give you all topics with the highest frequency in ambiguous cases, which can be useful if you need to handle edge scenarios.
内容的提问来源于stack exchange,提问作者Chris T.

