如何对data.table执行melt时仅对measure.vars首元素应用na.rm?
Solution: Filter Only
level_id NA Values Post-Melt The core issue here is that data.table::melt()'s na.rm = TRUE removes rows where any of the measure variables have NA values—including your level_date column, which you want to keep even if it's NA (as long as level_id is valid).
The cleanest, most efficient solution is to first perform the melt as you originally did, then filter out only the rows where level_id is NA using data.table's native row indexing. This avoids hacky workarounds like fake date values and keeps your code readable.
Step-by-Step Implementation
- First, recreate your original data.table (for reproducibility):
library(data.table) library(lubridate) dt.master <- data.table(user = seq(1,5), visit_id = c(2,4,NA,4,8), visit_date = c(dmy("10/02/2018"), dmy("11/04/2018"), NA, dmy("02/03/2018"), NA), offer_id = c(1,3,NA,NA,NA), offer_date = c(dmy("15/02/2018"), dmy("18/04/2018"), NA, NA, NA))
- Perform the melt, then chain a filter to remove rows where
level_idis NA:
dt.melted <- melt(dt.master, id.vars = "user", measure.vars = list(c("visit_id", "offer_id"), c("visit_date", "offer_date")), variable.name = "level", value.name = c("level_id", "level_date"))[!is.na(level_id)]
Result Verification
Running this code gives you exactly the output you want:
user level level_id level_date 1: 1 1 2 2018-02-10 2: 2 1 4 2018-04-11 3: 4 1 4 2018-03-02 4: 5 1 8 <NA> 5: 1 2 1 2018-02-15 6: 2 2 3 2018-04-18
Why This Works Better Than na.rm = TRUE
- The default
na.rm = TRUEinmelt()would drop the row for user 5 (level 1) becauselevel_dateis NA—this solution preserves it, since we only care aboutlevel_idbeing non-NA. - It's idiomatic data.table code: chaining operations keeps your logic concise and efficient, avoiding unnecessary intermediate objects.
- No messy workarounds (like filling fake dates) are needed, so your code stays maintainable.
内容的提问来源于stack exchange,提问作者illanxr
相关产品推荐
相关产品推荐

