You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

关于R语言数据聚合、去聚合及日期频次统计的技术咨询

Hey there! Let’s walk through your questions one by one, fix up the issues in your current code, and give you practical solutions for each task.

1. Fixing Your Aggregation & Data Clearing Code

First, let’s spot the small bugs in your existing code:

  • You accidentally included X4 twice in your cbind() call (duplicate entry)
  • When filtering ClearedData, you’re using mydata$X1 != 0 instead of the aggregated X1 from AggData—this will cause mismatched row indices!

Here’s the corrected version:

# Correct aggregation (removed duplicate X4)
AggData <- aggregate(cbind(X1, X2, X3, X4, X5, X6, X7) ~ CustomerNumber + Date + Accountnumber, 
                     data = mydata, 
                     FUN = sum)

# Filter out rows where aggregated X1 equals 0 (uses AggData's X1, not mydata's)
ClearedData <- AggData[AggData$X1 != 0, ]

2. Can You "Disaggregate" ClearedData?

Short answer: Only if you still have access to the original mydata dataset.

Aggregation is a one-way operation—once you sum multiple rows into a single aggregated row, you lose the individual values that made up that sum. Without the original data, there’s no way to split the aggregated totals back into their original records.

That said, if you do have mydata handy, you can get the original rows that correspond to the groups in ClearedData (which is probably what you mean by "disaggregating"). Here are two ways to do this:

Base R Method

# Keep only rows from mydata that match the groups in ClearedData
ClearedOriginalRows <- merge(mydata, 
                             ClearedData[, c("CustomerNumber", "Date", "Accountnumber")],
                             by = c("CustomerNumber", "Date", "Accountnumber"))

Tidyverse (dplyr) Method (Cleaner)

library(dplyr)

ClearedOriginalRows <- mydata %>%
  semi_join(ClearedData, by = c("CustomerNumber", "Date", "Accountnumber"))

This gives you the raw, unaggregated rows that belong to the groups where aggregated X1 wasn’t zero.

3. Adding a Date Frequency Variable

To count how often each date appears in ClearedData (or any dataset), you can use either base R or the tidyverse. Here are both approaches:

Base R Approach

# Calculate frequency of each date
date_counts <- table(ClearedData$Date)

# Add the frequency as a new column to ClearedData
ClearedData$DateFrequency <- as.numeric(date_counts[as.character(ClearedData$Date)])

Tidyverse (dplyr) Approach

This is more intuitive and handles grouping seamlessly:

library(dplyr)

ClearedData <- ClearedData %>%
  group_by(Date) %>%
  mutate(DateFrequency = n()) %>% # n() counts rows per group
  ungroup()

If you wanted to count frequency per CustomerNumber + Date instead, just adjust the group_by to group_by(CustomerNumber, Date).


内容的提问来源于stack exchange,提问作者Boetzka

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 04:06:06