You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言按q1a等多列分组统计观测数及频率并格式化输出问题

问题原因

你现有代码的问题是同时对q1a、q1a_30d、q1a_60d、q1a_90d四列联合分组,统计的是四列取值的组合频数,因此无法得到单列的统计结果。

修正方案

1. 统计代码

先将目标列转为长格式分组统计,再转回宽格式适配输出要求:

library(tidyverse)
library(formattable)
library(kableExtra)

stat_df <- df %>%
  # 选择四个目标列
  select(q1a, q1a_30d, q1a_60d, q1a_90d) %>%
  # 转为长格式,按时间段+响应值分组
  pivot_longer(everything(), names_to = "period", values_to = "response") %>%
  group_by(period, response) %>%
  summarise(cnt = n(), .groups = "drop") %>%
  # 计算各时间段内的频率
  group_by(period) %>%
  mutate(freq = percent(cnt / sum(cnt), 1)) %>%
  # 转回宽格式,每个时间段对应计数、频率两列
  pivot_wider(names_from = period, values_from = c(cnt, freq)) %>%
  # 调整列顺序匹配表头逻辑
  select(response,
         cnt_q1a, freq_q1a,
         cnt_q1a_30d, freq_q1a_30d,
         cnt_q1a_60d, freq_q1a_60d,
         cnt_q1a_90d, freq_q1a_90d) %>%
  # 保留响应值的原始因子排序
  mutate(response = factor(response,
                           levels = c("All of the time", "Very frequently", "Somewhat frequently", "Never", "No answer"))) %>%
  arrange(response)

2. 格式化输出代码

调整表头匹配列结构:

kable(stat_df,
      col.names = rep(c("计数", "频率"), 4),
      align = c("l", rep("cr", 4)),
      longtable = T,
      booktabs = T,
      valign = 't',
      escape = F,
      caption = '<b>佩戴口罩或面部遮盖物情况统计</b>') %>%
  kable_styling(bootstrap_options = c("striped", "hold_position"),
                full_width = T,
                position = "center",
                html_font = "Arial") %>%
  add_header_above(c("响应选项" = 1,
                     "基线" = 2,
                     "30天" = 2,
                     "60天" = 2,
                     "90天" = 2)) %>%
  column_spec(border_left = T, border_right = T)

如果你需要保留英文表头,将add_header_above和col.names里的中文替换为对应英文即可。


内容的提问来源于stack exchange,提问作者Meli

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.26 08:45:08