You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Stargazer与iwalk按组输出汇总统计时遇缺失值错误

问题:按分类变量输出汇总统计时触发stargazer错误

运行代码后持续收到如下错误:

Error in `map2()`:
ℹ 索引:1.
ℹ 名称:Control.
由`if (nchar(text.matrix[r, c]) > max.length[real.c]) ...`中的错误引起:
! 需要TRUE/FALSE的位置处有缺失值
回溯:
 1. ... %>% ...
 2. purrr::iwalk(...)
 3. purrr::walk2(.x, vec_index(.x), .f, ...)
 4. purrr::map2(.x, .y, .f, ..., .progress = .progress)
 5. purrr:::map2_("list", .x, .y, .f, ..., .progress = .progress)
 9. .f(.x[[i]], .y[[i]], ...)
10. stargazer::stargazer(...)
11. stargazer:::.stargazer.wrap(...)
12. stargazer (local) .text.output(latex.code)
13. stargazer (local) .text.column.width(t, c)

运行的代码如下:

df %>%
 split(. $treat) %>% 
 iwalk(~ 
     stargazer(., 
       type = "text",
       flip = TRUE,
       title = "Table X: Balance tests by treatment status",
       covariate.labels = c(
         "Female (0/1)",
         "Year of birth",
         "Phone number (0/1)",
         "Registered Democrat (0/1)",
         "Registered Republican (0/1)",
         "Unaffiliated (0/1)"),
        align = TRUE)
     )

已尝试的无效操作:

  • 移除标签中的(0/1)
  • 将treat变量转换为数值类型(原类型为因子)
  • 修改变量名移除下划线

可复现错误的数据片段:

df <- structure(list(treat = structure(c(1L, 1L, 1L, 1L, 1L, 2L, 2L, 
2L, 2L, 2L, 3L, 3L, 3L, 3L, 3L, 4L, 4L, 4L, 4L, 4L), levels = c("Control", 
"Treatment 1", "Treatment 2", "Treatment 3"
), class = "factor"), female = structure(c(1L, 1L, 2L, 2L, 2L, 
1L, 2L, 2L, 2L, 2L, 1L, 1L, 1L, 2L, 2L, 2L, 1L, 2L, 1L, 2L), levels = c("female", 
"male"), class = "factor"), birth_year = c(1945, 1930, 1990, 
1984, 1992, 1957, 1996, 1977, 1975, 1985, 1936, 1992, 1958, 1939, 
1986, 1955, 1962, 1973, 1986, 1950), provided_phone_no = c(1, 
1, 0, 1, 0, 1, 1, 0, 1, 0, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1), dem = c(1, 
1, 0, 0, 0, 0, 1, 1, 0, 1, 1, 1, 0, 1, 0, 1, 1, 1, 1, 0), rep = c(0, 
0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1), uaf = c(0, 
0, 1, 1, 0, 1, 0, 0, 1, 0, 0, 0, 1, 0, 1, 0, 0, 0, 0, 0)), class = c("grouped_df", 
"tbl_df", "tbl", "data.frame"), row.names = c(NA, -20L), groups = structure(list(
    treat = structure(1:4, levels = c("Control", "Treatment 1", 
    "Treatment 2", "Treatment 3"), class = "factor"), 
    .rows = structure(list(1:5, 6:10, 11:15, 16:20), ptype = integer(0), class = c("vctrs_list_of", 
    "vctrs_vctr", "list"))), row.names = c(NA, -4L), .drop = TRUE, class = c("tbl_df", 
"tbl", "data.frame")))

解决办法

问题根源:原数据框df是分组数据框(grouped_df),split()后每个子数据框仍保留分组属性,stargazer无法正确处理这类带分组标记的数据,从而触发缺失值判断错误。

修正步骤:

  1. 用ungroup()取消数据框的分组属性
  2. 确保covariate.labels与要展示的变量一一对应

修正后的代码:

df %>%
  ungroup() %>%  # 关键操作:取消分组属性
  split(.$treat) %>% 
  iwalk(~ 
          stargazer(., 
                    type = "text",
                    flip = TRUE,
                    title = paste0("Table X: Balance tests for ", .y),  # 可选:为每个分组设置专属标题
                    covariate.labels = c(
                      "Female (0/1)",
                      "Year of birth",
                      "Phone number (0/1)",
                      "Registered Democrat (0/1)",
                      "Registered Republican (0/1)",
                      "Unaffiliated (0/1)"),
                    align = TRUE)
        )

补充说明:

  • 取消分组后,split()生成的子数据框为普通数据框,stargazer可正常处理
  • 为每个分组单独设置标题,避免所有表格标题重复

内容的提问来源于stack exchange,提问作者C.Robin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 09:35:19