如何通过编程在PubChem中实现化合物名称的模糊搜索?
PubChem API 按名称获取相关化合物的解决方案
问题背景
在PubChem网页手动搜索时:
- 关键词
1-(2-Hydroxyphenyl)-2-phenyl ethanone:无完全匹配结果,但能找到4个部分匹配项; - 关键词
methane:显示1个最佳匹配(CID 297)+444027个相关化合物。
但使用以下代码调用API时,要么返回无匹配错误,要么仅返回最佳匹配结果:
import requests prolog = "https://pubchem.ncbi.nlm.nih.gov/rest/pug" a = requests.get(prolog + "/compound/name/1-(2-Hydroxyphenyl)-2-phenyl ethanone/json").json() print(a)
输出错误信息:
{'Fault': {'Code': 'PUGREST.NotFound', 'Message': 'No CID found', 'Details': ['No CID found that matches the given name']}}
原因分析
默认的/compound/name/接口仅返回精确匹配或优先级最高的最佳匹配化合物,当没有精确匹配时就会返回NotFound错误,这和网页搜索默认的模糊/相关匹配逻辑不一致。
解决方法
要获取相关匹配结果,可使用以下两种API接口:
1. 同义词搜索接口(/compound/name/synonyms/)
该接口会返回所有包含目标名称作为同义词、或名称相似的化合物,对应网页的部分匹配结果:
import requests prolog = "https://pubchem.ncbi.nlm.nih.gov/rest/pug" # 调用synonyms接口进行模糊名称搜索 response = requests.get(f"{prolog}/compound/name/synonyms/1-(2-Hydroxyphenyl)-2-phenyl ethanone/json") if response.status_code == 200: data = response.json() print(data) else: print(f"请求失败,状态码:{response.status_code}")
2. 通用搜索接口(/compound/search)
该接口支持更灵活的搜索配置,可通过参数控制返回结果数量、过滤条件等:
import requests prolog = "https://pubchem.ncbi.nlm.nih.gov/rest/pug" # 构造搜索参数,limit控制返回结果数量 params = { "q": "1-(2-Hydroxyphenyl)-2-phenyl ethanone", "format": "json", "limit": 10 # 可根据需求调整返回数量 } response = requests.get(f"{prolog}/compound/search", params=params) if response.status_code == 200: data = response.json() print(data) else: print(f"请求失败,状态码:{response.status_code}")
比如搜索methane时,使用这个接口并设置合适的limit值,就能获取到除最佳匹配外的其他相关化合物。
内容的提问来源于stack exchange,提问作者user19631495
相关产品推荐
相关产品推荐

