You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R中按ID分组生成序列并添加序列最大数值列

Solution for Adding Highest_number Column in R

Got it, let's break this down. You've already got the Sequence column working—resetting when the City repeats (within the same ID) or the ID changes. Now you need to add a Highest_number column that shows the maximum Sequence value for the consecutive City block each row belongs to, within its ID.

Using dplyr (Tidyverse)

First, let's start with your sample data, generate the Sequence column (included for full reproducibility), then add the Highest_number:

library(dplyr)

# Your sample data
mydf <- data.frame(
  ID = c(1, 1, 1, 1, 1, 1, 1, 1, 1, 2, 2, 2, 2, 2, 2, 2, 2),
  City = c("Nijmegen", "Nijmegen", "Arnhem", "Arnhem", "Nijmegen", "Arnhem", "Arnhem","Arnhem", "Nijmegen", "Nijmegen", "Utrecht", "Amsterdam", "Amsterdam", "Utrecht", "Utrecht", "Utrecht", "Utrecht")
)

# Step 1: Generate the Sequence column (as you've already implemented)
mydf <- mydf %>%
  group_by(ID) %>%
  mutate(
    # Create a helper column to identify consecutive City blocks
    city_block = cumsum(City != lag(City, default = first(City))),
    Sequence = ave(City, city_block, FUN = seq_along)
  ) %>%
  ungroup()

# Step 2: Add Highest_number by grouping on ID and city_block
mydf <- mydf %>%
  group_by(ID, city_block) %>%
  mutate(Highest_number = max(Sequence)) %>%
  ungroup() %>%
  # Remove the helper column if you don't need it
  select(-city_block)

Using data.table (Faster for Large Datasets)

If you're working with big data, data.table is more efficient. The rleid() function is perfect for identifying consecutive value runs:

library(data.table)

setDT(mydf)

# Generate both Sequence and Highest_number in one step
mydf[, `:=`(
  Sequence = seq_along(City),
  Highest_number = .N
), by = .(ID, rleid(City))]

How It Works

  • Consecutive Block Identification: Both methods first create a unique identifier for each consecutive run of the same City within an ID. In dplyr, we use cumsum(City != lag(City)); in data.table, rleid(City) does this in one step.
  • Calculate Maximum Sequence: For each block (grouped by ID and block identifier), we take the maximum value of Sequence (or just use .N in data.table, since Sequence starts at 1 and increments by 1—.N is the total rows in the block, which equals the max Sequence).

Result

Running either code will give you exactly the output you expected:

IDCitySequenceHighest_number
1Nijmegen12
1Nijmegen22
1Arnhem12
1Arnhem22
1Nijmegen11
1Arnhem13
1Arnhem23
1Arnhem33
1Nijmegen11
2Nijmegen11
2Utrecht11
2Amsterdam12
2Amsterdam22
2Utrecht14
2Utrecht24
2Utrecht34
2Utrecht44

内容的提问来源于stack exchange,提问作者Thijs

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 08:32:58