基于MATLAB从SDF文件精准提取指定ID分子的技术需求
MATLAB 提取SDF文件指定ID分子方案
准备工作
- 确认SDF文件中每个分子的ID有固定标识行(比如常见的
> <ID>,需匹配你的文件实际格式) - 目标ID文本文件需每行一个ID,无多余空格或空行
实现代码
1. 加载目标ID并构建查找集合
% 替换为你的ID文件路径 idList = readlines('target_ids.txt'); idList = strtrim(idList); % 用哈希集合加速ID查找 idSet = containers.Map(idList, ones(length(idList), 1));
2. 解析SDF并提取目标分子
% 替换为你的原始SDF路径和输出文件路径 inputFid = fopen('raw_molecules.sdf', 'r'); outputFid = fopen('extracted_molecules.sdf', 'w'); currentMol = []; currentId = ''; while ~feof(inputFid) line = fgetl(inputFid); if ischar(line) % 检测SDF分子结束符 if strcmp(line, '$$$$') currentMol = [currentMol, line, '\n']; % 匹配到目标ID则写入文件 if isKey(idSet, currentId) fprintf(outputFid, currentMol); remove(idSet, currentId); % 避免重复提取(可选) end % 重置缓存 currentMol = []; currentId = ''; else % 匹配ID标识行,这里根据你的SDF格式修改 if startsWith(line, '> <ID>') idLine = fgetl(inputFid); currentId = strtrim(idLine); currentMol = [currentMol, line, '\n', idLine, '\n']; else currentMol = [currentMol, line, '\n']; end end end end fclose(inputFid); fclose(outputFid);
3. 验证提取完整性
missingIds = keys(idSet); if ~isempty(missingIds) fprintf('未找到的ID:\n'); disp(missingIds); else fprintf('所有目标ID均已提取完成!\n'); end
关键注意事项
- 必须匹配SDF实际格式:如果你的SDF中ID标识不是
> <ID>,比如是> <MOLECULE_ID>,一定要修改代码中的判断条件 - 若SDF存在重复ID,
remove(idSet, currentId)会确保每个ID只提取一次,不需要的话可以注释掉 - 对于超大规模SDF(百万级分子),可以修改为分块读取,但数千级分子用上述代码完全稳定
内容的提问来源于stack exchange,提问作者Math Simp
相关产品推荐
相关产品推荐

