You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在data.table中重排列并输出结果?(无需setcolorder的链式实现方案)

Hey there! Let's figure out where your data.table code went wrong and fix it to match your dplyr workflow.

The Issue with Your Current Code

Your existing data.table code ends with [ , union(c("Amt_first", "Amt_last"), names(test_df)) ] — this line only returns a character vector of column names instead of rearranging the columns in the actual data.table. That's why you're seeing just the column names as output, not the full dataset with reordered columns.

Corrected Data.table Implementation

Here's the fixed code that replicates your dplyr result (adding the new columns, then moving them to the front without using setcolorder):

library(data.table)
library(lubridate)
library(tidyverse)

# Test dataset
test_df <- data.frame(id = c(1234, 1234, 5678, 5678), 
                      date = c("2021-10-10","2021-10-10", "2021-8-10", "2021-8-15"), 
                      Amount = c(54767, 96896, 34534, 79870)) %>% 
  mutate(date = ymd(date))

# Working data.table code
setDT(test_df)[order(date), 
               `:=`(Amt_first = first(Amount), Amt_last = last(Amount)), 
               by = id][, c("Amt_first", "Amt_last", setdiff(names(test_df), c("Amt_first", "Amt_last"))), with = FALSE]

Or a more concise version (using the .. shorthand for column selection, available in data.table 1.14.0+):

cols <- c("Amt_first", "Amt_last", setdiff(names(test_df), c("Amt_first", "Amt_last")))
setDT(test_df)[order(date), 
               `:=`(Amt_first = first(Amount), Amt_last = last(Amount)), 
               by = id][, ..cols]

Key Explanations

  1. Why setdiff instead of union?
    Your dplyr code uses everything() which preserves the original order of the remaining columns. union() would sort column names alphabetically, which breaks the original order. setdiff(names(test_df), c("Amt_first", "Amt_last")) keeps the non-new columns in their original sequence, just like everything().

  2. Why with = FALSE or ..cols?
    In data.table, passing a character vector directly to the j argument returns the vector itself. To tell data.table this vector represents column names to select, you either use with = FALSE or the .. prefix (which signals "look up this vector from the parent environment").

Expected Output

Running the corrected code will give you the exact same result as your dplyr implementation:

Amt_first Amt_last   id       date Amount
1:     54767    96896 1234 2021-10-10  54767
2:     54767    96896 1234 2021-10-10  96896
3:     34534    79870 5678 2021-08-10  34534
4:     34534    79870 5678 2021-08-15  79870

内容的提问来源于stack exchange,提问作者ViSa

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 20:57:27