如何使用Python基于RDFLib获取URIRef对应的本体字符串表示
解决方案
推荐优先使用RDFLib直接解析RDF语义数据的方案,比爬HTML更稳定,不会受页面样式调整影响。
方案1:基于RDFLib获取标注(推荐)
依赖安装
pip install rdflib requests
实现逻辑
OBO 本体的URIRef支持直接返回RDF格式的语义数据,我们只需在请求时指定对应格式的Accept头,拿到数据后用RDFLib提取rdfs:label属性即可,同时天然处理404无效链接的情况。
示例代码
import rdflib import requests from rdflib.namespace import RDFS from time import sleep def get_uriref_label(uri: str, delay: float = 0.5) -> str | None: # 避免请求频率过高 sleep(delay) headers = { "Accept": "text/turtle, application/rdf+xml" } try: resp = requests.get(uri, headers=headers, timeout=10) # 处理404等无效链接 if resp.status_code != 200: return None # 加载RDF数据 g = rdflib.Graph() g.parse(data=resp.text, format="turtle" if "text/turtle" in resp.headers.get("Content-Type", "") else "xml") # 提取rdfs:label值 for s, p, o in g.triples((rdflib.URIRef(uri), RDFS.label, None)): # 优先返回英文标注,需要其他语言可以调整语言过滤条件 if o.language is None or o.language == "en": return str(o) return None except Exception as e: print(f"处理{uri}时出错: {str(e)}") return None # 测试用例 uri_list = [ 'http://purl.obolibrary.org/obo/RO_0002219', 'http://purl.obolibrary.org/obo/RO_0002551', 'http://purl.obolibrary.org/obo/uberon/core#indirectly_supplies' ] result = {} for uri in uri_list: label = get_uriref_label(uri) result[uri] = label if label else "无有效标注" print(result)
运行后输出示例:
{ 'http://purl.obolibrary.org/obo/RO_0002219': 'surrounded by', 'http://purl.obolibrary.org/obo/RO_0002551': 'has skeleton part', 'http://purl.obolibrary.org/obo/uberon/core#indirectly_supplies': '无有效标注' }
方案2:解析HTML页面获取红色文本(不推荐)
如果确实需要按页面样式提取红色文本,可以用BeautifulSoup解析页面,该方案稳定性依赖OBO官网页面结构不变。
依赖安装
pip install requests beautifulsoup4
示例代码
import requests from bs4 import BeautifulSoup from time import sleep def get_label_from_html(uri: str, delay: float = 0.5) -> str | None: sleep(delay) try: resp = requests.get(uri, timeout=10) if resp.status_code != 200: return None soup = BeautifulSoup(resp.text, "html.parser") # 匹配红色样式的span标签 red_span = soup.find("span", style=lambda value: value and "color: red" in value) return red_span.get_text(strip=True) if red_span else None except Exception as e: print(f"处理{uri}时出错: {str(e)}") return None
批量处理注意事项
- 可根据需求调整请求间隔
delay参数,避免请求频率过高被限制访问 - 建议添加本地缓存逻辑,已处理过的URIRef无需重复请求,大幅提升批量处理效率
内容的提问来源于stack exchange,提问作者Penguin
相关产品推荐
相关产品推荐

