You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R语言中基于多变量生成结果统计表的技术问询

Great question! When dealing with multiple columns of the same type (like your PR1 to PR25 surgery codes), the key is to first reshape your data from wide to long format—this makes it easy to aggregate counts across all variables. Here's how you can do this in R using tidyverse tools, which are perfect for this kind of task:

Step 1: Reshape Wide Data to Long Format

First, we'll convert your wide data frame into a long format where each row represents a single surgery code entry from any PR column:

# Load required library
library(tidyverse)

# Replace "your_data" with your actual data frame name
long_data <- your_data %>%
  pivot_longer(
    cols = starts_with("PR"),  # Target all columns starting with "PR" (works for PR1-PR25)
    names_to = "PR_variable",  # New column to store which PR column the code came from
    values_to = "surgery_code" # New column to store the surgery code itself
  )

Step 2: Generate the Summary Table

Now we can calculate key stats like total frequency, which PR columns each code appears in, and percentages:

summary_table <- long_data %>%
  group_by(surgery_code) %>%
  summarize(
    total_count = n(),  # Total number of times the code appears across all PR columns
    unique_PR_columns = n_distinct(PR_variable),  # How many different PR columns include this code
    PR_columns_list = paste(unique(PR_variable), collapse = ", "),  # List of specific PR columns
    percentage_of_total = round((total_count / nrow(long_data)) * 100, 2)  # Percentage of all entries
  ) %>%
  arrange(desc(total_count))  # Sort from most to least frequent codes

Example Output (Using Your Sample Data)

For your provided sample data, the resulting summary_table would look like this:

surgery_codetotal_countunique_PR_columnsPR_columns_listpercentage_of_total
22233PR1, PR2, PR320.00
52722PR1, PR213.33
56923PR1, PR2, PR313.33
34122PR1, PR313.33
142222PR2, PR313.33
160011PR16.67
166011PR36.67

Base R Alternative (No Tidyverse)

If you prefer not to use tidyverse, here's a base R approach:

# Reshape using stack()
long_data_base <- stack(your_data[, grep("^PR", colnames(your_data))])

# Calculate frequency counts
count_table <- as.data.frame(table(long_data_base$values))
colnames(count_table) <- c("surgery_code", "total_count")

# Add percentage
count_table$percentage_of_total <- round((count_table$total_count / nrow(long_data_base)) * 100, 2)

# Sort by frequency
count_table <- count_table[order(-count_table$total_count), ]

Both methods scale seamlessly to your full PR1-PR25 dataset—no need to list each column individually!

内容的提问来源于stack exchange,提问作者TimF

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 08:28:53