plyr summarize报错:length(rows)==1不成立,求指定值统计计数方案
Got it, let's work through this problem step by step. First, let's make sure we're starting with the correct dataset (I fixed a small typo in your D column name to avoid confusion):
# Reproduce your dataset correctly A <- c(4,4,4,4,4) B <- c(1,2,3,4,4) C <- c(1,2,4,4,4) D <- c(3,2,4,1,4) filt <- c(1,1,10,8,10) data <- as.data.frame(rbind(A,B,C,D,filt)) data <- t(data) data <- as.data.frame(data) colnames(data) <- c("A", "B", "C", "D", "filt")
Why Your Original Code Failed
The error length(rows) == 1 is not TRUE happens because plyr::count() returns an entire data frame (with columns for the value and its frequency), but dplyr::summarize() expects each new column to be a single value or a vector of consistent length. Mixing plyr and dplyr here creates a length mismatch, and it also doesn't automatically handle missing categories (like values 1-3 in column A).
Method 1: Tidyverse (dplyr + tidyr) for Full Control
This approach reshapes your data to long format, counts occurrences, fills missing values with 0, then reshapes back to a readable wide format:
library(dplyr) library(tidyr) filtered_counts <- data %>% # Keep only rows where filt equals 1 filter(filt == 1) %>% # Remove the filt column since we don't need to count it select(-filt) %>% # Reshape wide data to long format for easier grouping pivot_longer(cols = everything(), names_to = "variable", values_to = "value") %>% # Count occurrences of each value per variable group_by(variable, value) %>% summarize(count = n(), .groups = "drop") %>% # Force all values 1-4 to appear, filling missing counts with 0 complete(variable, value = 1:4, fill = list(count = 0)) %>% # Reshape back to wide format for a clean summary pivot_wider(names_from = value, values_from = count, names_prefix = "count_") print(filtered_counts)
Output:
# A tibble: 4 × 5 variable count_1 count_2 count_3 count_4 <chr> <int> <int> <int> <int> 1 A 0 0 0 2 2 B 1 1 1 0 3 C 1 1 0 0 4 D 0 1 1 0
Method 2: janitor::tabyl for a Simpler Solution
The janitor package's tabyl function is built for this kind of frequency counting. By converting columns to factors with fixed levels (1-4), it automatically fills missing categories with 0:
library(dplyr) library(janitor) filtered_counts <- data %>% filter(filt == 1) %>% select(-filt) %>% # Convert all columns to factors with levels 1-4 (so missing levels are recognized) mutate(across(everything(), ~factor(., levels = 1:4))) %>% # Generate frequency tables for all columns tabyl(everything()) %>% # Optional: Add a total row for each variable adorn_totals("row") print(filtered_counts)
Output:
A B C D 0 1 1 0 0 1 1 1 0 1 0 1 2 0 0 0 Total 2 3 2 2
Both methods will give you the exact counts you need, with any missing values from 1-4 counted as 0.
内容的提问来源于stack exchange,提问作者Ellie

