R语言中统计矩阵元素频次及生成计数矩阵的技术咨询
Hey there! Let's walk through how to build your Mat matrix step by step, which will count the frequency of each unique element in every row of your original dataset.
Step 1: Recreate Your Dataset
First, let's make sure we have your data frame set up properly in R. Based on the sample you provided, here's how to define it (just add the missing 3 rows to each column vector if needed):
# Define your original dataset data <- data.frame( var1 = c("A", "B", "D", "D"), var2 = c("C", "B", "D", "B"), var3 = c("A", "A", "B", "B"), var4 = c("C", "C", "B", "C") )
Step 2: Get Unique Elements (Your rule Vector)
Instead of manually typing out the unique elements, we can extract them automatically from your data to avoid errors. If you already know the exact set of elements (A, B, C, D), you can also define rule directly—both methods work!
# Option 1: Auto-extract unique elements (sorted for consistency) rule <- sort(unique(unlist(data))) # Option 2: Manually define if you prefer # rule <- c("A", "B", "C", "D")
Step 3: Build the Frequency Matrix Mat
We'll create a helper function to count how many times each element in rule appears in a single row, then apply this function to every row of your dataset. Finally, we'll transpose the result to get the correct row/column structure:
# Helper function to count element frequencies for one row count_row_frequencies <- function(row) { # Use factor to ensure all elements in rule are included (even if count is 0) table(factor(row, levels = rule)) } # Apply the function to each row, transpose to get rows matching original data Mat <- t(apply(data, 1, count_row_frequencies)) # Add column names for clarity colnames(Mat) <- rule
Step 4: Check the Result
If you print Mat, you'll see a matrix where each row corresponds to a row in your original data, and each column shows the count of that element:
# View the final matrix Mat
Sample output (for the 4 rows you provided):
A B C D [1,] 2 0 2 0 [2,] 1 2 1 0 [3,] 0 2 0 2 [4,] 0 2 1 1
This approach works no matter how many rows your original data has—just make sure the dataset is defined correctly, and the code will handle the rest!
内容的提问来源于stack exchange,提问作者user3642360

