使用Stargazer与iwalk按组输出汇总统计时遇缺失值错误
问题:按分类变量输出汇总统计时触发stargazer错误
运行代码后持续收到如下错误:
Error in `map2()`: ℹ 索引:1. ℹ 名称:Control. 由`if (nchar(text.matrix[r, c]) > max.length[real.c]) ...`中的错误引起: ! 需要TRUE/FALSE的位置处有缺失值 回溯: 1. ... %>% ... 2. purrr::iwalk(...) 3. purrr::walk2(.x, vec_index(.x), .f, ...) 4. purrr::map2(.x, .y, .f, ..., .progress = .progress) 5. purrr:::map2_("list", .x, .y, .f, ..., .progress = .progress) 9. .f(.x[[i]], .y[[i]], ...) 10. stargazer::stargazer(...) 11. stargazer:::.stargazer.wrap(...) 12. stargazer (local) .text.output(latex.code) 13. stargazer (local) .text.column.width(t, c)
运行的代码如下:
df %>% split(. $treat) %>% iwalk(~ stargazer(., type = "text", flip = TRUE, title = "Table X: Balance tests by treatment status", covariate.labels = c( "Female (0/1)", "Year of birth", "Phone number (0/1)", "Registered Democrat (0/1)", "Registered Republican (0/1)", "Unaffiliated (0/1)"), align = TRUE) )
已尝试的无效操作:
- 移除标签中的
(0/1) - 将
treat变量转换为数值类型(原类型为因子) - 修改变量名移除下划线
可复现错误的数据片段:
df <- structure(list(treat = structure(c(1L, 1L, 1L, 1L, 1L, 2L, 2L, 2L, 2L, 2L, 3L, 3L, 3L, 3L, 3L, 4L, 4L, 4L, 4L, 4L), levels = c("Control", "Treatment 1", "Treatment 2", "Treatment 3" ), class = "factor"), female = structure(c(1L, 1L, 2L, 2L, 2L, 1L, 2L, 2L, 2L, 2L, 1L, 1L, 1L, 2L, 2L, 2L, 1L, 2L, 1L, 2L), levels = c("female", "male"), class = "factor"), birth_year = c(1945, 1930, 1990, 1984, 1992, 1957, 1996, 1977, 1975, 1985, 1936, 1992, 1958, 1939, 1986, 1955, 1962, 1973, 1986, 1950), provided_phone_no = c(1, 1, 0, 1, 0, 1, 1, 0, 1, 0, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1), dem = c(1, 1, 0, 0, 0, 0, 1, 1, 0, 1, 1, 1, 0, 1, 0, 1, 1, 1, 1, 0), rep = c(0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1), uaf = c(0, 0, 1, 1, 0, 1, 0, 0, 1, 0, 0, 0, 1, 0, 1, 0, 0, 0, 0, 0)), class = c("grouped_df", "tbl_df", "tbl", "data.frame"), row.names = c(NA, -20L), groups = structure(list( treat = structure(1:4, levels = c("Control", "Treatment 1", "Treatment 2", "Treatment 3"), class = "factor"), .rows = structure(list(1:5, 6:10, 11:15, 16:20), ptype = integer(0), class = c("vctrs_list_of", "vctrs_vctr", "list"))), row.names = c(NA, -4L), .drop = TRUE, class = c("tbl_df", "tbl", "data.frame")))
解决办法
问题根源:原数据框df是分组数据框(grouped_df),split()后每个子数据框仍保留分组属性,stargazer无法正确处理这类带分组标记的数据,从而触发缺失值判断错误。
修正步骤:
- 用
ungroup()取消数据框的分组属性 - 确保
covariate.labels与要展示的变量一一对应
修正后的代码:
df %>% ungroup() %>% # 关键操作:取消分组属性 split(.$treat) %>% iwalk(~ stargazer(., type = "text", flip = TRUE, title = paste0("Table X: Balance tests for ", .y), # 可选:为每个分组设置专属标题 covariate.labels = c( "Female (0/1)", "Year of birth", "Phone number (0/1)", "Registered Democrat (0/1)", "Registered Republican (0/1)", "Unaffiliated (0/1)"), align = TRUE) )
补充说明:
- 取消分组后,
split()生成的子数据框为普通数据框,stargazer可正常处理 - 为每个分组单独设置标题,避免所有表格标题重复
内容的提问来源于stack exchange,提问作者C.Robin
相关产品推荐
相关产品推荐

