R语言技术问询:按DateTime分组统计Category列不含“ 1”后缀的唯一值计数
Solution for Counting Non-" 1" Ending Categories by DateTime
Got it, I see what you need here—let's build that df2 step by step. The core of the problem is filtering out the Category values that end with " 1", then counting the remaining entries per DateTime and Category pair.
Option 1: Using Tidyverse (dplyr)
This is the most readable approach if you're working with the tidy ecosystem:
First, load the required package (if you haven't already):
library(dplyr)
Then run this pipeline to create df2:
df2 <- df %>% # Filter out any Category that ends with " 1" (using regex to match the end of the string) filter(!grepl(" 1$", Category)) %>% # Group by both DateTime and Category to count per unique pair group_by(DateTime, Category) %>% # Calculate the count and drop grouping to get a flat data frame summarise(CatCount = n(), .groups = "drop") %>% # Reorder columns to match your expected output select(Category, DateTime, CatCount)
Breakdown of each step:
!grepl(" 1$", Category): The regex1$precisely matches strings that end with " 1". The!negates the match, so we keep only rows where Category doesn't end with " 1".group_by(DateTime, Category): Groups the data so we can count entries for each unique combination of time and valid category.summarise(CatCount = n()): Counts the number of rows in each group, naming the count columnCatCount..groups = "drop"ensures we don't keep the grouping structure after summarizing.select(...): Reorders columns to match your desired output order (Category first, then DateTime, then CatCount).
Option 2: Base R (No External Packages)
If you prefer to stick with base R without loading tidyverse, here's how to do it:
# First, filter the data to exclude Category entries ending with " 1" filtered_df <- df[!grepl(" 1$", df$Category), ] # Use table() to count occurrences per DateTime and Category pair count_table <- table(filtered_df$DateTime, filtered_df$Category) # Convert the table to a data frame and rename columns df2 <- as.data.frame(count_table, responseName = "CatCount") colnames(df2) <- c("DateTime", "Category", "CatCount") # Reorder columns to match your expected output df2 <- df2[, c("Category", "DateTime", "CatCount")]
Both approaches will produce exactly the output you requested:
| Category | DateTime | CatCount |
|---|---|---|
| A | 2022-08-29 00:00:00 | 2 |
| B | 2022-08-29 00:00:00 | 3 |
| A | 2022-08-29 02:00:00 | 1 |
| B | 2022-08-29 02:00:00 | 3 |
内容的提问来源于stack exchange,提问作者Jacob
相关产品推荐
相关产品推荐

