data.table中j运算符使用困惑:如何按组计算矩阵归一化值
Hey there! I totally get how the j operator in data.table can feel a bit opaque at first—let's work through this to get your normalization formula applied in a way that fits your existing workflow.
First, let's recap your current code for context:
library(data.table) DT = data.table(Grp1=letters[1:10], Grp2=letters[rep(1:5, 2)], Grp3=letters[rep(1:2, 5)], matrix(1:100, 10)) grp2.cluster.mean = DT[, lapply(.SD, mean), by=Grp2, .SDcols=-c("Grp1", "Grp3")] variable.means = DT[, lapply(.SD, mean), .SDcols=-c("Grp1", "Grp2", "Grp3")]
Your goal is to normalize the numeric matrix columns using the formula (cluster mean - variable mean) / (cluster mean + variable mean), where cluster means are grouped by Grp2, and variable means are the global averages across all rows. Here are two approaches that stick to data.table's syntax:
Approach 1: Normalize cluster-level means (matches your grp2.cluster.mean structure)
This method takes your precomputed mean tables, combines them, and applies the formula directly:
# Convert global means to a long-format table for easy joining variable.means_long = melt(as.data.table(t(variable.means)), variable.name = "var", value.name = "global_mean") # Convert cluster means to long format, join with global means, then calculate normalization normalized_cluster_means = melt(grp2.cluster.mean, id.vars = "Grp2", variable.name = "var", value.name = "cluster_mean")[, normalized_value := (cluster_mean - global_mean)/(cluster_mean + global_mean), on = .(var), global_mean = variable.means_long$global_mean ] # Optional: Convert back to wide format to match your original cluster mean table structure normalized_cluster_wide = dcast(normalized_cluster_means, Grp2 ~ var, value.var = "normalized_value")
Approach 2: Add normalized columns directly to the original DT
If you want to keep the original row-level data alongside the normalized values (grouped by Grp2), this works:
# Store global means as a named vector for easy reference global_means_vec = unlist(variable.means) # Define a helper function for your normalization formula to keep code clean normalize_val = function(cluster_mean, global_mean) { (cluster_mean - global_mean)/(cluster_mean + global_mean) } # Group by Grp2, compute normalized values for each numeric column DT_with_normalized = DT[, c(.(Grp1, Grp2, Grp3), lapply(names(global_means_vec), function(col) { normalize_val(mean(get(col)), global_means_vec[[col]]) })), by = Grp2] # Rename the new normalized columns for clarity setnames(DT_with_normalized, paste0("normalized_", names(global_means_vec)), 4:ncol(DT_with_normalized))
Both approaches lean into data.table's core [i, j, by] syntax:
- In Approach 1, we use
melt/dcastto reshape data for joining, then apply the formula in thejslot with a join condition. - In Approach 2, we use
lapplyinjto iterate over numeric columns, compute the cluster mean per group, and apply the helper function with the precomputed global mean.
内容的提问来源于stack exchange,提问作者CharlesM

