You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R语言中筛选出与癌症症状相关的产品清单

问题:找出与癌症症状相关的风险产品清单

我有一份FDA食品、膳食补充剂及化妆品的消费者投诉数据集,已完成数据清洗,并通过以下R代码构建了症状0-1矩阵:

syms <- strsplit(dat$symptoms, ", ")
tm   <- matrix(0, nrow=nrow(dat), ncol=length(unique(unlist(syms))))
colnames(tm) <- unique(unlist(syms))

for(i in 1:length(syms)) {
  tm[i, syms[[i]]] <- 1
}
dat$symptoms <- NULL

其中dat包含患者投诉核心字段:

receivedid...productoutcome
9/30/20222022-CFS-014640...centrum silver men's 50+other outcome
9/30/20222022-CFS-014637...liquid collagen shotlife threatening

tm为症状0-1矩阵,示例结构如下:

diarrhoeavomitingcancer
010
100
......1

我需要找出关联癌症症状、需注意的产品清单,尝试了以下代码但未成功:

# Find rows in tm matrix where the "cancer" symptom is present
cancer_rows <- which(tm[, "cancer"] == 1)

# Create a vector of product names associated with "cancer" symptoms
products_to_avoid <- unique(dat$product[cancer_rows])

请问如何正确实现该需求?


解决方案

1. 先排查核心问题:症状列名匹配

首先确认tm矩阵中是否存在准确名为"cancer"的列——常见问题是症状名存在大小写(如"Cancer")、拼写变体(如"carcinoma")或空格,导致索引失败。可以用colnames(tm)查看所有症状列名,确认目标列的准确名称。

2. 基础修正代码

如果列名确实是"cancer",可以简化代码(无需which(),直接用逻辑向量索引更稳健):

# 直接筛选出tm中cancer列为1的行对应的产品
products_to_avoid <- unique(dat$product[tm[, "cancer"] == 1])

3. 稳健适配版(处理列名不匹配)

如果不确定症状列名的准确写法,用模糊匹配定位目标列:

# 不区分大小写查找包含"cancer"关键词的列
cancer_col <- grep("cancer", colnames(tm), ignore.case = TRUE)

if(length(cancer_col) > 0){
  # 提取对应产品
  products_to_avoid <- unique(dat$product[tm[, cancer_col] == 1])
} else {
  warning("未找到与癌症相关的症状列")
  products_to_avoid <- character(0)
}

4. 合并数据框的直观方式(避免索引错位)

将症状矩阵合并回原数据框,再筛选,彻底避免行索引不匹配的问题:

# 将症状矩阵转为数据框并合并到原dat
dat_sym <- cbind(dat, as.data.frame(tm))
# 筛选有癌症症状的行,提取唯一产品
products_to_avoid <- unique(dat_sym$product[dat_sym$cancer == 1])

注意事项

上述结果仅为消费者投诉中关联癌症症状的产品清单,并非经过临床验证的致癌产品,仅作为初步风险参考。

内容的提问来源于stack exchange,提问作者cinnamond

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 15:38:13