You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言中pairwise_survdiff无法识别数据列的问题求助

解决pairwise_survdiff无法识别带空格列名的问题

问题背景

使用survminer包的pairwise_survdiff()进行生存分析时,出现列名无法识别的错误,但这些带空格的列名在coxph()、survfit()等函数中均可正常调用。

首次尝试代码及报错

pairwise_survdiff(
  formula = Surv(`Days in Treatment - Interuptions`, Outcome == "Cured") ~ 
    `Reoccuring Case` + 
    `Number of Times previously Cured` + 
    `Number of Times Previously Defaulted` + 
    `IYCF Counselling Received` + 
    `Medication Teaching Completed` + 
    `Estimated Days of missed treatment`,
  data = OTPData_V6,
  p.adjust.method = "bonferroni", 
  na.action = na.omit
)

报错信息:

Error in `vectbl_as_col_location()`:
! Can't subset columns that don't exist.
✖ Columns `\`Reoccuring Case\``, `\`Number of Times previously Cured\``, `\`Number of Times Previously Defaulted\``, `\`IYCF Counselling Received\``, `\`Medication Teaching Completed\``, etc. don't exist.
Run `rlang::last_error()` to see where the error occurred.

二次尝试(直接引用数据框列)及报错

改用OTPData_V6$列名``方式调用,仍报错:

pairwise_survdiff(formula=Surv(`Days in Treatment - Interuptions`,Outcome=="Cured")~ 
                    OTPData_V6$`Reoccuring Case`+ 
                    OTPData_V6$`Number of Times previously Cured`+ OTPData_V6$`Number of Times Previously Defaulted`+ 
                    OTPData_V6$`IYCF Counselling Received`+ OTPData_V6$`Medication Teaching Completed`
                  + OTPData_V6$`Estimated Days of missed treatment`,
                  data=OTPData_V6, p.adjust.method = "bonferroni",na.action = na.omit)

报错信息:

Error in `vectbl_as_col_location()`:
! Can't subset columns that don't exist.
✖ Columns `OTPData_V6$\`Reoccuring Case\``, `OTPData_V6$\`Number of Times previously Cured\``, `OTPData_V6$\`Number of Times Previously Defaulted\``, `OTPData_V6$\`IYCF Counselling Received\``, `OTPData_V6$\`Medication Teaching Completed\``, etc. don't exist.
Run `rlang::last_error()` to see where the error occurred.

数据集头部信息

dput(head(OTPData_V6, 2))
structure(list(`OTP Registration Number` = c(70, 65), `MTI Unique Patient ID` = c(70, 
65), Address = c("Medebay Zana", "Setit Humera"), Nationality = c("Ethiopia", 
"Ethiopia"), Region = c("Tigray", "Tigray"), Zone = c("Northwestern Zone", 
"Western Tigray"), Woreda = c("Medebay Zana", "Kafta Humera"), 
    `Age at first Admission (Months)` = c(18, 9), `Age at Discharge` = c(23, 
    13), `Age recorded at OTP Registration (Months)` = c(18, 
    9), Sex = structure(c(1L, 1L), levels = c("Female", "Male"
    ), class = "factor"), `Reoccuring Case` = structure(c(1L, 
    1L), levels = c("No", "Yes"), class = "factor"), `New Admission` = structure(c(1L, 
    1L), levels = c("No", "Yes"), class = "factor"), `Transfer or Readmission` = structure(c(2L, 
    2L), levels = c("No", "Yes"), class = "factor"), `Number of Times Previously Admitted to OTP` = c(0, 
    0), `Number of Times previously Cured` = c(0, 0), `Number of Times Previously Defaulted` = c(0, 
    0), `Admission Weight (KG)` = c(5, 6.1), `Admission MUAC` = c(9.4, 
    10.2), `Discharge Date` = c("11/5/2021", "11/5/2021"), `Length of Treatment Period (Days)` = c(154, 
    135), `Days in Treatment - Interuptions` = c(140, 121), `Total Time Elapsed from First OTP Admission` = c(154, 
    135), `Total Days in Treatment since first admission` = c(140, 
    121), `Discharge Weight (KG)` = c(NA_real_, NA_real_), `Discharge MUAC` = c(NA_real_, 
    NA_real_), Outcome = structure(c(3L, 3L), levels = c("Cured", 
    "Default", "Default due to stockout", "Ongoing"), class = "factor"), 
    `Number of Technical Defaults during treatment period` = c(1, 
    1), `Number of stockout periods during treatment` = c(2, 
    2), `Estimated Days of missed treatment` = c(14, 14), `Number of stockout periods during treatment - Warehouse` = c(2, 
    2), `Estimated Days of missed treatment - Warehouse` = c(14, 
    14), Displaced = structure(c(2L, 2L), levels = c("Host", 
    "Yes"), class = "factor"), `First OTP Registration Date` = c("6/4/2021", 
    "6/23/2021"), `OTP registration date` = c("6/4/2021", "6/23/2021"
    ), `Arrival Date` = c("6/4/2021", "6/23/2021"), NCD = structure(c(1L, 
    1L), levels = c("No", "Yes"), class = "factor"), `Communicable (Infectious) Dx` = structure(c(1L, 
    1L), levels = c("No", "Yes"), class = "factor"), Trauma = c(0, 
    0), `Acute Bloody Diarrhea` = structure(c(1L, 1L), levels = c("No", 
    "Yes"), class = "factor"), `Acute Respiratory Infection` = structure(c(1L, 
    1L), levels = c("No", "Yes"), class = "factor"), `Acute Watery Diarrhea` = structure(c(1L, 
    1L), levels = c("No", "Yes"), class = "factor"), `Diarrhea Unspecified` = structure(c(1L, 
    1L), levels = c("No", "Yes"), class = "factor"), Dysentery = c(0, 
    0), `Fever or unknown origin` = c(0, 0), Pneumonia = structure(c(1L, 
    1L), levels = c("No", "Yes"), class = "factor"), Scabies = structure(c(1L, 
    1L), levels = c("No", "Yes"), class = "factor"), `Suspected Cholera` = c(0, 
    0), `Suspected Covid-19` = c(0, 0), `Suspected Hepatitis` = c(0, 
    0), `Suspected HIV` = c(0, 0), `Suspected Intestinal Parasite` = structure(c(1L, 
    1L), levels = c("No", "Yes"), class = "factor"), `Suspected Malaria` = c(0, 
    0), `Suspected Measles` = c(0, 0), `Suspected Meningitis` = c(0, 
    0), `Suspected TB` = c(0, 0), `Skin Infection` = structure(c(1L, 
    1L), levels = c("No", "Yes"), class = "factor"), `Typhoid Fever` = c(0, 
    0), `Urinary Tract Infection` = c(0, 0), Conjunctivitis = structure(c(1L, 
    1L), levels = c("No", "Yes"), class = "factor"), `IYCF Counselling Received` = structure(c(1L, 
    1L), levels = c("No", "Yes"), class = "factor"), `Prescriptions Filled: Yes` = structure(c(1L, 
    1L), levels = c("No", "Yes"), class = "factor"), `Prescriptions Filled: No` = structure(c(1L, 
    1L), levels = c("No", "Yes"), class = "factor"), `Medication Teaching Completed` = structure(c(1L, 
    1L), levels = c("No", "Yes"), class = "factor"), `Referral Made` = c(0, 
    0)), row.names = c(NA, -2L), class = c("tbl_df", "tbl", "data.frame"
))

解决方案

该问题源于pairwise_survdiff()对带空格列名的转义解析异常,以下三种方法可解决:

方法1:重命名列名(推荐)

将带空格和特殊字符的列名替换为下划线分隔的格式,彻底避免解析问题:

# 批量替换列名中的空格和短横线
colnames(OTPData_V6) <- gsub("[ -]", "_", colnames(OTPData_V6))

# 调用函数
pairwise_survdiff(
  formula = Surv(Days_in_Treatment__Interuptions, Outcome == "Cured") ~ 
    Reoccuring_Case + 
    Number_of_Times_previously_Cured + 
    Number_of_Times_Previously_Defaulted + 
    IYCF_Counselling_Received + 
    Medication_Teaching_Completed + 
    Estimated_Days_of_missed_treatment,
  data = OTPData_V6,
  p.adjust.method = "bonferroni", 
  na.action = na.omit
)

方法2:用reformulate()构建公式

通过函数生成公式,规避手动转义错误:

# 定义协变量列名向量
covars <- c(
  "`Reoccuring Case`",
  "`Number of Times previously Cured`",
  "`Number of Times Previously Defaulted`",
  "`IYCF Counselling Received`",
  "`Medication Teaching Completed`",
  "`Estimated Days of missed treatment`"
)

# 生成公式
formula <- reformulate(
  termlabels = covars,
  response = quote(Surv(`Days in Treatment - Interuptions`, Outcome == "Cured"))
)

# 调用函数
pairwise_survdiff(
  formula = formula,
  data = OTPData_V6,
  p.adjust.method = "bonferroni", 
  na.action = na.omit
)

方法3:转换为普通data.frame

tibble格式的列名解析逻辑可能与普通data.frame有差异,转换后尝试:

# 转换为普通data.frame
OTPData_V6_df <- as.data.frame(OTPData_V6)

# 调用函数
pairwise_survdiff(
  formula = Surv(`Days in Treatment - Interuptions`, Outcome == "Cured") ~ 
    `Reoccuring Case` + 
    `Number of Times previously Cured` + 
    `Number of Times Previously Defaulted` + 
    `IYCF Counselling Received` + 
    `Medication Teaching Completed` + 
    `Estimated Days of missed treatment`,
  data = OTPData_V6_df,
  p.adjust.method = "bonferroni", 
  na.action = na.omit
)

内容的提问来源于stack exchange,提问作者Jared

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 19:42:00