You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Structural Notation Type拆分结构值至SMILES/InChI列的问题

按类型拆分结构表示到对应列的解决方案

需求说明

需要将Structural Notation列的内容,依据Structural Notation Type列的标记,分别填入新建的SMILES和InChI列:SMILES类型的值放入SMILES列,INCHI类型的值放入InChI列。

原尝试问题

此前通过创建Combine列,为SMILES行添加&&分隔符、INCHI行添加%%分隔符,再用data.table的tstrsplit函数拆分,结果出现错位(如第3行SMILES值被分到InChI列)。


方法1:无需额外列的直接赋值

直接利用data.table的条件赋值逻辑,不需要创建中间列:

library(data.table)
# 假设数据表为dt
dt[, SMILES := ifelse(`Structural Notation Type` == "SMILES", `Structural Notation`, NA)]
dt[, InChI := ifelse(`Structural Notation Type` == "INCHI", `Structural Notation`, NA)]

如果存在重复分组(比如同一实体对应两种结构类型),可按分组聚合赋值:

dt[, c("SMILES", "InChI") := .(
  first(ifelse(`Structural Notation Type` == "SMILES", `Structural Notation`, NA)),
  first(ifelse(`Structural Notation Type` == "INCHI", `Structural Notation`, NA))
), by = .(你的分组列名)]

方法2:修正tstrsplit的错误用法

原方法错误在于使用了两种不同分隔符,tstrsplit需要固定分隔符才能正确拆分。统一分隔符后再处理:

# 创建带统一分隔符的Combine列
dt[, Combine := paste(`Structural Notation Type`, `Structural Notation`, sep = "|")]
# 拆分后按类型匹配赋值
split_result <- tstrsplit(dt$Combine, "|", fixed = TRUE)
dt[split_result[[1]] == "SMILES", SMILES := split_result[[2]]]
dt[split_result[[1]] == "INCHI", InChI := split_result[[2]]]
# 清理中间列
dt[, Combine := NULL]

此方法可行但冗余,推荐用直接赋值或dcast方法。

方法3:推荐的dcast解决方案(理想结果)

使用data.table的dcast函数直接转置行到列,是最简洁的实现方式:

# 生成临时行ID(若无分组列时用)
dt[, id := .I]
# 转置类型列为SMILES/InChI列
result_dt <- dcast(dt, id ~ `Structural Notation Type`, value.var = "Structural Notation")
# 调整列名(按需)
setnames(result_dt, c("SMILES", "INCHI"), c("SMILES", "InChI"))
# 移除临时ID列
result_dt[, id := NULL]

该方法直接生成每行对应SMILES和InChI值的结果,不会出现错位问题。


内容的提问来源于stack exchange,提问作者iembry

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.11 00:33:17