寻求高效工具:7000条植物文本的Plant Ontology等本体批量注释
高效植物本体注释工具与代码方案
工具方案
- OBO Foundry 本体映射工具:直接对接Plant Ontology(PO)、Crop Ontology(CO)、Food Ontology等官方本体库,支持批量上传文本列表。可通过配置过滤规则,指定仅匹配目标本体,避免NCBI术语干扰,精准输出预期注释结果。
- OntoMaton:支持自定义导入OBO格式的本体文件,将PO、CO、Food Ontology导入后,可针对植物文本列表做批量匹配。能调整匹配阈值,减少无关结果,提升注释精准度。
- NCBI BioPortal 高级模式:之前使用Annotator+未得到预期结果,大概率是未配置数据源过滤。进入高级设置,将PO、CO、Food Ontology设为唯一匹配数据源,关闭NCBI相关本体的匹配权限,即可得到目标本体的注释。
代码方案
Python + OWLready2 批量匹配
OWLready2可直接加载OWL/OBO格式的本体文件,通过简单的字符串匹配逻辑实现批量注释,示例代码如下:
from owlready2 import * # 加载目标本体 po = get_ontology("http://purl.obolibrary.org/obo/po.owl").load() co = get_ontology("http://purl.obolibrary.org/obo/co.owl").load() food_ont = get_ontology("http://purl.obolibrary.org/obo/foodon.owl").load() # 读取植物文本列表(假设每行一个植物名称) with open("plant_list.txt", "r", encoding="utf-8") as f: plant_names = [line.strip() for line in f if line.strip()] annotation_results = [] for name in plant_names: matches = [] # 匹配Plant Ontology的类名与同义词 for cls in po.classes(): label = str(cls.label[0]) if cls.label else "" synonyms = getattr(cls, 'synonym', []) synonym_strs = [str(syn) for syn in synonyms] if name.lower() in label.lower() or any(name.lower() in s.lower() for s in synonym_strs): matches.append(f"PO: {label} ({cls.iri})") # 匹配Crop Ontology for cls in co.classes(): label = str(cls.label[0]) if cls.label else "" synonyms = getattr(cls, 'synonym', []) synonym_strs = [str(syn) for syn in synonyms] if name.lower() in label.lower() or any(name.lower() in s.lower() for s in synonym_strs): matches.append(f"CO: {label} ({cls.iri})") # 匹配Food Ontology for cls in food_ont.classes(): label = str(cls.label[0]) if cls.label else "" synonyms = getattr(cls, 'synonym', []) synonym_strs = [str(syn) for syn in synonyms] if name.lower() in label.lower() or any(name.lower() in s.lower() for s in synonym_strs): matches.append(f"FoodON: {label} ({cls.iri})") annotation_results.append({name: matches}) # 保存注释结果 with open("annotation_results.txt", "w", encoding="utf-8") as f: for res in annotation_results: plant, annots = next(iter(res.items())) f.write(f"{plant}: {', '.join(annots) if annots else '无匹配结果'}\n")
SPARQL 批量查询
若目标本体托管在SPARQL端点(如EBI Ontology Lookup Service),可编写批量SPARQL查询,指定目标本体命名空间,精准检索匹配术语。示例查询逻辑:
PREFIX po: <http://purl.obolibrary.org/obo/PO_> PREFIX co: <http://purl.obolibrary.org/obo/CO_> PREFIX foodon: <http://purl.obolibrary.org/obo/FOODON_> PREFIX rdfs: <http://www.w3.org/2000/01/rdf-schema#> PREFIX oboInOwl: <http://www.geneontology.org/formats/oboInOwl#> SELECT ?term ?label ?synonym WHERE { VALUES ?query {"植物名称1" "植物名称2" ...} { ?term rdfs:label ?label . FILTER(LCASE(?label) = LCASE(?query)) } UNION { ?term oboInOwl:hasExactSynonym ?synonym . FILTER(LCASE(?synonym) = LCASE(?query)) } FILTER(STRSTARTS(STR(?term), "http://purl.obolibrary.org/obo/PO_") || STRSTARTS(STR(?term), "http://purl.obolibrary.org/obo/CO_") || STRSTARTS(STR(?term), "http://purl.obolibrary.org/obo/FOODON_")) }
可通过Python的SPARQLWrapper库实现批量查询,自动替换?query中的植物名称,批量获取注释结果。
内容的提问来源于stack exchange,提问作者Agnes
相关产品推荐
相关产品推荐

