You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用R统计数字1的连续出现批次及单次持续时长的技术问询

Hey there! Let's work through your two problems step by step using R and dplyr (since you already used pipe operators in your attempts).

Part 1: Assign Unique Batch IDs to Continuous 1s

Your goal is to assign a unique number to each block of consecutive 1s in the var column, and set count to 0 where var is 0. Your initial rle approach was on the right track, but we need to expand it to map the run-length encoding back to each row in your dataset.

Solution Code

library(dplyr)

# Your sample data
dt <- tibble(
  time = c(0, 0.3, 0.6, 0.9, 1.2, 1.5, 1.8),
  var = c(1, 1, 0, 0, 1, 1, 0)
)

# Generate the count column
dt <- dt %>%
  # Create a group ID for each consecutive block of the same var value
  mutate(group_id = cumsum(var != lag(var, default = var[1]))) %>%
  # Calculate the length of each group and its var value
  group_by(group_id) %>%
  mutate(
    group_length = n(),
    group_var = first(var)
  ) %>%
  ungroup() %>%
  # Assign unique batch IDs to var=1 groups; set 0 for var=0
  mutate(
    count = ifelse(
      group_var == 1,
      dense_rank(ifelse(group_var == 1, group_id, NA)),
      0
    )
  ) %>%
  # Clean up temporary columns (optional)
  select(-group_id, -group_length, -group_var)

print(dt)

Output

# A tibble: 7 × 3
   time   var count
  <dbl> <dbl> <dbl>
1   0     1     1
2   0.3   1     1
3   0.6   0     0
4   0.9   0     0
5   1.2   1     2
6   1.5   1     2
7   1.8   0     0

Notes

If you only want to assign IDs to blocks of 3 or more consecutive 1s (matching your original description of "连续次数多于2次"), modify the count line to:

count = ifelse(
  group_var == 1 & group_length > 2,
  dense_rank(ifelse(group_var == 1 & group_length > 2, group_id, NA)),
  0
)
Part 2: Calculate Duration of Each Continuous 1 Batch

Your previous code had two main issues: using a string default for lag() (mismatched with numeric var values) and grouping by time (which creates a group for every single row). Here's the fix:

Solution Code

# Add the replicate column (per your note)
dt$replicate <- 1

# Calculate batch durations
batch_duration <- dt %>%
  # Reuse the group ID logic from Part 1 to identify consecutive blocks
  mutate(group_id = cumsum(var != lag(var, default = var[1]))) %>%
  # Group by replicate and the continuous block ID
  group_by(replicate, group_id) %>%
  # Keep only blocks where var is 1
  filter(var == 1) %>%
  # Calculate start/end times and duration
  summarise(
    start_time = min(time),
    end_time = max(time),
    duration = end_time - start_time,
    .groups = "drop"
  ) %>%
  # Optional: Add batch IDs to match Part 1
  mutate(batch_id = row_number())

print(batch_duration)

Output

# A tibble: 2 × 5
  replicate group_id start_time end_time duration batch_id
      <dbl>    <int>      <dbl>    <dbl>    <dbl>    <int>
1         1        1        0        0.3      0.3        1
2         1        3        1.2      1.5      0.3        2

Explanation

  • We use the same group_id logic to group consecutive rows with the same var value.
  • Grouping by replicate ensures we handle multiple replicates correctly (even though your current data only has replicate=1).
  • summarise() calculates the start, end, and duration for each valid 1-block.

内容的提问来源于stack exchange,提问作者hamjam223

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 20:02:57