R语言中pairwise_survdiff无法识别数据列的问题求助
解决pairwise_survdiff无法识别带空格列名的问题
问题背景
使用survminer包的pairwise_survdiff()进行生存分析时,出现列名无法识别的错误,但这些带空格的列名在coxph()、survfit()等函数中均可正常调用。
首次尝试代码及报错
pairwise_survdiff( formula = Surv(`Days in Treatment - Interuptions`, Outcome == "Cured") ~ `Reoccuring Case` + `Number of Times previously Cured` + `Number of Times Previously Defaulted` + `IYCF Counselling Received` + `Medication Teaching Completed` + `Estimated Days of missed treatment`, data = OTPData_V6, p.adjust.method = "bonferroni", na.action = na.omit )
报错信息:
Error in `vectbl_as_col_location()`: ! Can't subset columns that don't exist. ✖ Columns `\`Reoccuring Case\``, `\`Number of Times previously Cured\``, `\`Number of Times Previously Defaulted\``, `\`IYCF Counselling Received\``, `\`Medication Teaching Completed\``, etc. don't exist. Run `rlang::last_error()` to see where the error occurred.
二次尝试(直接引用数据框列)及报错
改用OTPData_V6$列名``方式调用,仍报错:
pairwise_survdiff(formula=Surv(`Days in Treatment - Interuptions`,Outcome=="Cured")~ OTPData_V6$`Reoccuring Case`+ OTPData_V6$`Number of Times previously Cured`+ OTPData_V6$`Number of Times Previously Defaulted`+ OTPData_V6$`IYCF Counselling Received`+ OTPData_V6$`Medication Teaching Completed` + OTPData_V6$`Estimated Days of missed treatment`, data=OTPData_V6, p.adjust.method = "bonferroni",na.action = na.omit)
报错信息:
Error in `vectbl_as_col_location()`: ! Can't subset columns that don't exist. ✖ Columns `OTPData_V6$\`Reoccuring Case\``, `OTPData_V6$\`Number of Times previously Cured\``, `OTPData_V6$\`Number of Times Previously Defaulted\``, `OTPData_V6$\`IYCF Counselling Received\``, `OTPData_V6$\`Medication Teaching Completed\``, etc. don't exist. Run `rlang::last_error()` to see where the error occurred.
数据集头部信息
dput(head(OTPData_V6, 2)) structure(list(`OTP Registration Number` = c(70, 65), `MTI Unique Patient ID` = c(70, 65), Address = c("Medebay Zana", "Setit Humera"), Nationality = c("Ethiopia", "Ethiopia"), Region = c("Tigray", "Tigray"), Zone = c("Northwestern Zone", "Western Tigray"), Woreda = c("Medebay Zana", "Kafta Humera"), `Age at first Admission (Months)` = c(18, 9), `Age at Discharge` = c(23, 13), `Age recorded at OTP Registration (Months)` = c(18, 9), Sex = structure(c(1L, 1L), levels = c("Female", "Male" ), class = "factor"), `Reoccuring Case` = structure(c(1L, 1L), levels = c("No", "Yes"), class = "factor"), `New Admission` = structure(c(1L, 1L), levels = c("No", "Yes"), class = "factor"), `Transfer or Readmission` = structure(c(2L, 2L), levels = c("No", "Yes"), class = "factor"), `Number of Times Previously Admitted to OTP` = c(0, 0), `Number of Times previously Cured` = c(0, 0), `Number of Times Previously Defaulted` = c(0, 0), `Admission Weight (KG)` = c(5, 6.1), `Admission MUAC` = c(9.4, 10.2), `Discharge Date` = c("11/5/2021", "11/5/2021"), `Length of Treatment Period (Days)` = c(154, 135), `Days in Treatment - Interuptions` = c(140, 121), `Total Time Elapsed from First OTP Admission` = c(154, 135), `Total Days in Treatment since first admission` = c(140, 121), `Discharge Weight (KG)` = c(NA_real_, NA_real_), `Discharge MUAC` = c(NA_real_, NA_real_), Outcome = structure(c(3L, 3L), levels = c("Cured", "Default", "Default due to stockout", "Ongoing"), class = "factor"), `Number of Technical Defaults during treatment period` = c(1, 1), `Number of stockout periods during treatment` = c(2, 2), `Estimated Days of missed treatment` = c(14, 14), `Number of stockout periods during treatment - Warehouse` = c(2, 2), `Estimated Days of missed treatment - Warehouse` = c(14, 14), Displaced = structure(c(2L, 2L), levels = c("Host", "Yes"), class = "factor"), `First OTP Registration Date` = c("6/4/2021", "6/23/2021"), `OTP registration date` = c("6/4/2021", "6/23/2021" ), `Arrival Date` = c("6/4/2021", "6/23/2021"), NCD = structure(c(1L, 1L), levels = c("No", "Yes"), class = "factor"), `Communicable (Infectious) Dx` = structure(c(1L, 1L), levels = c("No", "Yes"), class = "factor"), Trauma = c(0, 0), `Acute Bloody Diarrhea` = structure(c(1L, 1L), levels = c("No", "Yes"), class = "factor"), `Acute Respiratory Infection` = structure(c(1L, 1L), levels = c("No", "Yes"), class = "factor"), `Acute Watery Diarrhea` = structure(c(1L, 1L), levels = c("No", "Yes"), class = "factor"), `Diarrhea Unspecified` = structure(c(1L, 1L), levels = c("No", "Yes"), class = "factor"), Dysentery = c(0, 0), `Fever or unknown origin` = c(0, 0), Pneumonia = structure(c(1L, 1L), levels = c("No", "Yes"), class = "factor"), Scabies = structure(c(1L, 1L), levels = c("No", "Yes"), class = "factor"), `Suspected Cholera` = c(0, 0), `Suspected Covid-19` = c(0, 0), `Suspected Hepatitis` = c(0, 0), `Suspected HIV` = c(0, 0), `Suspected Intestinal Parasite` = structure(c(1L, 1L), levels = c("No", "Yes"), class = "factor"), `Suspected Malaria` = c(0, 0), `Suspected Measles` = c(0, 0), `Suspected Meningitis` = c(0, 0), `Suspected TB` = c(0, 0), `Skin Infection` = structure(c(1L, 1L), levels = c("No", "Yes"), class = "factor"), `Typhoid Fever` = c(0, 0), `Urinary Tract Infection` = c(0, 0), Conjunctivitis = structure(c(1L, 1L), levels = c("No", "Yes"), class = "factor"), `IYCF Counselling Received` = structure(c(1L, 1L), levels = c("No", "Yes"), class = "factor"), `Prescriptions Filled: Yes` = structure(c(1L, 1L), levels = c("No", "Yes"), class = "factor"), `Prescriptions Filled: No` = structure(c(1L, 1L), levels = c("No", "Yes"), class = "factor"), `Medication Teaching Completed` = structure(c(1L, 1L), levels = c("No", "Yes"), class = "factor"), `Referral Made` = c(0, 0)), row.names = c(NA, -2L), class = c("tbl_df", "tbl", "data.frame" ))
解决方案
该问题源于pairwise_survdiff()对带空格列名的转义解析异常,以下三种方法可解决:
方法1:重命名列名(推荐)
将带空格和特殊字符的列名替换为下划线分隔的格式,彻底避免解析问题:
# 批量替换列名中的空格和短横线 colnames(OTPData_V6) <- gsub("[ -]", "_", colnames(OTPData_V6)) # 调用函数 pairwise_survdiff( formula = Surv(Days_in_Treatment__Interuptions, Outcome == "Cured") ~ Reoccuring_Case + Number_of_Times_previously_Cured + Number_of_Times_Previously_Defaulted + IYCF_Counselling_Received + Medication_Teaching_Completed + Estimated_Days_of_missed_treatment, data = OTPData_V6, p.adjust.method = "bonferroni", na.action = na.omit )
方法2:用reformulate()构建公式
通过函数生成公式,规避手动转义错误:
# 定义协变量列名向量 covars <- c( "`Reoccuring Case`", "`Number of Times previously Cured`", "`Number of Times Previously Defaulted`", "`IYCF Counselling Received`", "`Medication Teaching Completed`", "`Estimated Days of missed treatment`" ) # 生成公式 formula <- reformulate( termlabels = covars, response = quote(Surv(`Days in Treatment - Interuptions`, Outcome == "Cured")) ) # 调用函数 pairwise_survdiff( formula = formula, data = OTPData_V6, p.adjust.method = "bonferroni", na.action = na.omit )
方法3:转换为普通data.frame
tibble格式的列名解析逻辑可能与普通data.frame有差异,转换后尝试:
# 转换为普通data.frame OTPData_V6_df <- as.data.frame(OTPData_V6) # 调用函数 pairwise_survdiff( formula = Surv(`Days in Treatment - Interuptions`, Outcome == "Cured") ~ `Reoccuring Case` + `Number of Times previously Cured` + `Number of Times Previously Defaulted` + `IYCF Counselling Received` + `Medication Teaching Completed` + `Estimated Days of missed treatment`, data = OTPData_V6_df, p.adjust.method = "bonferroni", na.action = na.omit )
内容的提问来源于stack exchange,提问作者Jared
相关产品推荐
相关产品推荐

