如何判断元素所属列表?R语言data.table列表名称匹配实现
Solution
First, let's recap your setup to ensure we're working with the same data:
# Define your category list catlist <- list( c("< 30 days","1-3 months","4-6 months"), c("7-12 months"), c("1-3 years"), c("4-5 years","5+ years","never") ) names(catlist) <- 1:length(catlist) # Load data.table and create your temp table library(data.table) tmp <- data.table(variable = c("never","1-3 years"))
1. Add new_variable to map values to list names
The most efficient way (especially for large datasets) is to create a lookup table first, then join it to your tmp data.table. This leverages data.table's fast join operations:
# Create a lookup table that maps each category to its list name lookup <- data.table( variable = unlist(catlist), new_variable = rep(names(catlist), sapply(catlist, length)) ) # Join the lookup table to tmp to add the new column tmp <- merge(tmp, lookup, by = "variable", all.x = TRUE)
Alternatively, you can use an update join to modify tmp in place without creating a new object:
tmp[lookup, on = "variable", new_variable := i.new_variable]
After either approach, your tmp table will look like this:
variable new_variable 1: 1-3 years 3 2: never 4
2. Check if elements are present in catlist
To verify if values in variable exist anywhere in catlist, use the %in% operator with unlist(catlist) (which flattens the list into a single vector):
tmp[, is_present := variable %in% unlist(catlist)]
This adds a logical column where TRUE means the value is found in catlist, and FALSE means it's not. For your sample data, both rows will show TRUE.
Alternative: Using sapply for small datasets
If you're working with a small dataset and prefer a more concise (though less efficient) approach, you can use nested sapply calls to directly map each value to its list name:
tmp[, new_variable := sapply(variable, function(x) { names(catlist)[sapply(catlist, function(y) x %in% y)] })]
This works by checking each value against every element in catlist and returning the corresponding name. However, for large datasets, the lookup table method is significantly faster.
内容的提问来源于stack exchange,提问作者quant

