You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R中绘制频率与序列时长的关系——时间戳转换问题

解决R中绘制流程步骤随归一化时间频率的问题

1. 数据预处理:归一化每个流程的时间

先将每个流程的时间锚定到自身起始点(即该流程第一个步骤的开始时间),计算每个步骤的相对开始/结束时间(以秒为单位,相对于流程起始点),这样就能统一不同流程的时间基准。

library(tidyverse)
library(lubridate)

# 替换为你的实际数据框
df <- tibble(
  ProcessID = c("X", "X", "Y", "Y", "Y"),
  Step = c("a", "d", "a", "c", "d"),
  Start = dmy_hm(c("30.09.2022 14:08", "30.09.2022 14:11", "29.09.2022 11:55", "29.09.2022 13:16", "29.09.2022 14:26")),
  End = dmy_hm(c("30.09.2022 14:11", "30.09.2022 14:24", "29.09.2022 13:16", "29.09.2022 14:26", "30.09.2022 17:17")),
  Seconds = c(165, 756, 4876, 4199, 96654)
)

# 计算每个流程的起始时间,再生成相对时间
df_normalized <- df %>%
  group_by(ProcessID) %>%
  mutate(
    process_start = min(Start),
    rel_start = as.numeric(difftime(Start, process_start, units = "secs")),
    rel_end = as.numeric(difftime(End, process_start, units = "secs"))
  ) %>%
  ungroup()

2. 生成时间序列点

确定观察的时间范围(从0到所有流程中最长的相对结束时间),生成均匀间隔的时间点(示例按60秒/分钟间隔,可按需调整):

# 获取所有流程的最长相对结束时间
max_time <- max(df_normalized$rel_end)

# 生成时间点序列
time_points <- seq(from = 0, to = max_time, by = 60)

3. 统计每个时间点的步骤频率

对每个时间点,统计有多少个流程当前处于该步骤:

# 构建时间点与步骤的匹配表,判断活跃状态并统计频率
frequency_data <- expand_grid(
  time = time_points,
  step_row = 1:nrow(df_normalized)
) %>%
  left_join(df_normalized, by = c("step_row" = "row_number()")) %>%
  mutate(is_active = time >= rel_start & time < rel_end) %>%
  group_by(time, Step) %>%
  summarise(frequency = sum(is_active), .groups = "drop")

4. 绘制时间轴频率图

用ggplot2绘制折线图或面积图,直观展示各步骤随时间的频率变化:

# 折线图版本
ggplot(frequency_data, aes(x = time, y = frequency, color = Step)) +
  geom_line(linewidth = 1) +
  labs(
    x = "流程归一化时间(秒)",
    y = "活跃流程数量",
    title = "各步骤随流程时间的频率变化",
    color = "步骤"
  ) +
  theme_minimal()

# 面积图版本(更直观展示占比)
# ggplot(frequency_data, aes(x = time, y = frequency, fill = Step)) +
#   geom_area(alpha = 0.7, position = "stack") +
#   labs(
#     x = "流程归一化时间(秒)",
#     y = "活跃流程数量",
#     title = "各步骤随流程时间的频率变化",
#     fill = "步骤"
#   ) +
#   theme_minimal()

补充提示

  • 若数据量较大(1000个流程),可增大时间点间隔(如300秒/5分钟)减少计算量。
  • 若步骤a在单个流程中多次出现,代码会自动识别每个a的时间段,统计该时间点处于任意a步骤的流程数量。
  • 可通过scale_x_continuous(labels = function(x) x/3600)将x轴转换为小时单位,提升可读性。

内容的提问来源于stack exchange,提问作者b7616

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 02:46:07