基于R生成的数据框统计特定动物和时间下的同病农户数量
Hey there! Let's work through how to solve your problem of counting the number of farmers with the same disease, filtered to a specific animal and time point.
First, let's recap your dataset setup code (I added a seed for reproducibility so we can all test with the same random data):
# Define categorical vectors farmer <- c("F1","F2","F3","F4","F5","F6") animal <- c("A1","A2","A3","A4","A5","A6","A7","A8") time <- c("T1","T2","T3","T4","T5") disease <- c("ww","ee","dd","ss","kk","ll","xx","oo","hh") # Generate the random dataset set.seed(123) # Ensures consistent random results df <- data.frame( farmer = sample(farmer, size=100, replace = TRUE), animal = sample(animal, size=100, replace = TRUE), time = sample(time, size=100, replace = TRUE), disease = sample(disease, size=100, replace = TRUE) ) # Preview the first few rows head(df)
Let's say we want to target animal "A1" and time "T1" (you can swap these values for any specific animal/time you need). Here are two straightforward ways to get the count of farmers per disease:
Method 1: Using Base R
We'll filter the data first, then aggregate to count unique farmers per disease:
# Define your target animal and time target_animal <- "A1" target_time <- "T1" # Filter data and count unique farmers per disease result_base <- aggregate( farmer ~ disease, data = subset(df, animal == target_animal & time == target_time), FUN = function(x) length(unique(x)) ) # Rename the count column for clarity colnames(result_base)[2] <- "farmer_count" # View the result print(result_base)
Explanation:
subset(df, animal == target_animal & time == target_time)narrows down the dataset to only rows matching your specific animal and time.aggregate(farmer ~ disease, ..., FUN = function(x) length(unique(x)))groups the filtered data by disease, then counts how many unique farmers are in each group. Usingunique()ensures we don't count the same farmer multiple times if they have duplicate records for the same animal/time/disease.
Method 2: Using dplyr (Tidyverse)
If you prefer a more readable, pipe-based approach, dplyr is perfect for this:
First, make sure you have the tidyverse installed and loaded:
# Install if needed # install.packages("tidyverse") library(tidyverse)
Then run this code:
# Define target values target_animal <- "A1" target_time <- "T1" # Calculate farmer count per disease result_dplyr <- df %>% filter(animal == target_animal, time == target_time) %>% group_by(disease) %>% summarize(farmer_count = n_distinct(farmer)) %>% ungroup() # View the result print(result_dplyr)
Explanation:
filter()narrows down the data to your target animal and time.group_by(disease)groups the filtered rows by each disease.summarize(farmer_count = n_distinct(farmer))calculates the number of unique farmers in each disease group.ungroup()removes the grouping structure, making the result easier to work with later.
Quick Notes:
- If you want to count total records (instead of unique farmers), replace
length(unique(x))withlength(x)in base R, orn_distinct(farmer)withn()in dplyr. - You can easily update
target_animalandtarget_timeto any values from youranimalandtimevectors to get results for other combinations.
内容的提问来源于stack exchange,提问作者Martini Sun

