如何在data.table中重排列并输出结果?(无需setcolorder的链式实现方案)
Hey there! Let's figure out where your data.table code went wrong and fix it to match your dplyr workflow.
The Issue with Your Current Code
Your existing data.table code ends with [ , union(c("Amt_first", "Amt_last"), names(test_df)) ] — this line only returns a character vector of column names instead of rearranging the columns in the actual data.table. That's why you're seeing just the column names as output, not the full dataset with reordered columns.
Corrected Data.table Implementation
Here's the fixed code that replicates your dplyr result (adding the new columns, then moving them to the front without using setcolorder):
library(data.table) library(lubridate) library(tidyverse) # Test dataset test_df <- data.frame(id = c(1234, 1234, 5678, 5678), date = c("2021-10-10","2021-10-10", "2021-8-10", "2021-8-15"), Amount = c(54767, 96896, 34534, 79870)) %>% mutate(date = ymd(date)) # Working data.table code setDT(test_df)[order(date), `:=`(Amt_first = first(Amount), Amt_last = last(Amount)), by = id][, c("Amt_first", "Amt_last", setdiff(names(test_df), c("Amt_first", "Amt_last"))), with = FALSE]
Or a more concise version (using the .. shorthand for column selection, available in data.table 1.14.0+):
cols <- c("Amt_first", "Amt_last", setdiff(names(test_df), c("Amt_first", "Amt_last"))) setDT(test_df)[order(date), `:=`(Amt_first = first(Amount), Amt_last = last(Amount)), by = id][, ..cols]
Key Explanations
Why
setdiffinstead ofunion?
Your dplyr code useseverything()which preserves the original order of the remaining columns.union()would sort column names alphabetically, which breaks the original order.setdiff(names(test_df), c("Amt_first", "Amt_last"))keeps the non-new columns in their original sequence, just likeeverything().Why
with = FALSEor..cols?
In data.table, passing a character vector directly to thejargument returns the vector itself. To tell data.table this vector represents column names to select, you either usewith = FALSEor the..prefix (which signals "look up this vector from the parent environment").
Expected Output
Running the corrected code will give you the exact same result as your dplyr implementation:
Amt_first Amt_last id date Amount 1: 54767 96896 1234 2021-10-10 54767 2: 54767 96896 1234 2021-10-10 96896 3: 34534 79870 5678 2021-08-10 34534 4: 34534 79870 5678 2021-08-15 79870
内容的提问来源于stack exchange,提问作者ViSa

