在R中按列统计层级计数及矩阵等位基因列频的实现方法
Hey there! Let's break down your two R questions step by step—first the general case of counting levels per column, then the specific allele matrix scenario.
There are a few solid approaches depending on whether you prefer base R or the tidyverse ecosystem. Let's start with a sample dataset to demo:
# Create a sample data frame with categorical columns set.seed(123) sample_df <- data.frame( Group = sample(c("Control", "Treatment"), 15, replace = TRUE), Status = sample(c("Positive", "Negative", "Neutral"), 15, replace = TRUE), Category = sample(c("X", "Y"), 15, replace = TRUE) )
Option 1: Base R with apply() + table()
This is quick and doesn't require any extra packages:
# Count levels for each column (returns a list of tables) column_level_counts <- apply(sample_df, 2, table) # View the result column_level_counts
The output is a list where each element corresponds to a column, showing the count of each level present.
Option 2: Tidyverse (dplyr + tidyr) for a structured data frame
If you want the results in a clean, tabular format (easier to work with for further analysis), use this workflow:
library(tidyverse) sample_df %>% # Convert wide data to long format pivot_longer(cols = everything(), names_to = "Column", values_to = "Level") %>% # Count occurrences of each level per column count(Column, Level) %>% # Convert back to wide format, filling missing levels with 0 pivot_wider(names_from = Level, values_from = n, values_fill = 0)
This gives you a data frame where each row is an original column, and each column is a level with its count.
Option 3: Janitor package for simplified tabulation
The janitor package has a handy tabyl() function that makes this even easier:
library(janitor) library(purrr) # Generate a tabyl for each column and combine into a single data frame map(sample_df, tabyl) %>% map_dfr(~ as.data.frame(.x), .id = "Column")
This includes both counts and percentages, which can be useful for quick summaries.
For your specific case with an A/T/C/G allele matrix, we can adapt the above methods to ensure we always include all four alleles (even if one isn't present in a column). Let's start with a sample allele matrix:
# Simulate an allele matrix (rows = samples, columns = loci) set.seed(456) allele_matrix <- matrix( sample(c("A", "T", "C", "G"), 50, replace = TRUE), nrow = 10, ncol = 5, dimnames = list(paste0("Sample", 1:10), paste0("Locus", 1:5)) )
Method 1: Tidyverse Workflow (Recommended for Readability)
library(tidyverse) allele_counts <- allele_matrix %>% # Convert matrix to data frame as.data.frame() %>% # Reshape to long format pivot_longer(cols = everything(), names_to = "Locus", values_to = "Allele") %>% # Count each allele per locus count(Locus, Allele) %>% # Ensure all four alleles are included (fill missing with 0) pivot_wider( names_from = Allele, values_from = n, values_fill = 0, names_expand = TRUE # Forces inclusion of all possible alleles ) # View the final dataset allele_counts
This produces a clean data frame where each row is a locus (original column), and columns show the count of A, T, C, and G.
Method 2: Base R for No Extra Dependencies
If you prefer sticking to base R, use this approach to force all four alleles to appear:
# Define the four possible alleles alleles <- c("A", "T", "C", "G") # Count alleles per column, ensuring all four are included count_list <- apply(allele_matrix, 2, function(col) { table(factor(col, levels = alleles)) }) # Convert the list of tables to a data frame allele_counts_df <- as.data.frame(t(count_list)) # View the result allele_counts_df
The factor() function ensures that even if an allele is missing from a column, it still shows up with a count of 0.
内容的提问来源于stack exchange,提问作者user11709754

