RDFLib+JSON-LD+命名图:使用@graph时如何覆盖图名称?
如何强制RDFLib把JSON-LD所有数据塞进指定命名图?
最近折腾JSON-LD和RDFLib的时候碰到个小麻烦:我要处理一个包含博客文章信息和类定义的JSON-LD文档,本来想把所有数据都导入到https://myblog.net/rdf/这个命名图里,结果RDFLib按规范拆成了两个图,这可不是我想要的效果。先给大家理理整个场景:
我的JSON-LD文档
我用JSON-LD描述了一篇博客文章,同时用@graph关键字定义了Robot类和它的子类Rover,文档内容是这样的:
{ "@context": { "schema": "https://schema.org/", "blog": "https://blog.me/", "ex": "https://example.org/", "rdfs": "http://www.w3.org/2000/01/rdf-schema#" }, "@id": "blog:JSONLD-and-named-graphs", "@type": "schema:blogPost", "rdfs:label": "JSON-LD and Named Graphs", "@graph": [ { "@id": "ex:Robot", "@type": "rdfs:Class" }, { "@id": "ex:Rover", "rdfs:subClassOf": { "@id": "ex:Robot" } } ] }
最初的RDFLib代码
我写了下面的代码,想把所有数据导入到指定的命名图里:
from rdflib import ConjunctiveGraph import json # 假设JSONLD_DOCUMENT就是上面的JSON对象 graph = ConjunctiveGraph() serialized_document = json.dumps(JSONLD_DOCUMENT) graph.parse( data=serialized_document, format='json-ld', publicID='https://myblog.net/rdf/', )
预期vs实际结果
- 我想要的结果:所有数据(包括博客文章的属性和
@graph里的类定义)都乖乖待在https://myblog.net/rdf/这个命名图里 - 实际发生的事:生成了两个命名图:
https://myblog.net/rdf/:只有博客文章的ID、类型和标签数据https://blog.me/JSONLD-and-named-graphs:装着@graph里的所有类定义
后来才明白,这根本不是bug,是JSON-LD规范的要求——{"@id": "xxx", "@graph": {...}}这种结构本身就是在声明一个名为@id值的命名图,所以RDFLib会严格按规范拆分数据。
那怎么绕过规范,强制把所有数据塞进指定命名图?
给大家分享两个可行的方法:
方法1:预处理JSON-LD,扁平化结构
最简单的方式就是修改JSON-LD文档,把顶层的博客文章属性也放进@graph数组里,让整个文档变成一个统一的@graph结构,这样解析时就不会生成额外的命名图了。修改后的文档如下:
{ "@context": { "schema": "https://schema.org/", "blog": "https://blog.me/", "ex": "https://example.org/", "rdfs": "http://www.w3.org/2000/01/rdf-schema#" }, "@graph": [ { "@id": "blog:JSONLD-and-named-graphs", "@type": "schema:blogPost", "rdfs:label": "JSON-LD and Named Graphs" }, { "@id": "ex:Robot", "@type": "rdfs:Class" }, { "@id": "ex:Rover", "rdfs:subClassOf": { "@id": "ex:Robot" } } ] }
再用原来的代码解析,所有数据就都会导入到你指定的publicID命名图里了。
方法2:解析后手动合并命名图
如果不想修改原始的JSON-LD文档,也可以先按正常流程解析,然后把额外生成的命名图里的三元组全部移到目标图,再删掉那个多余的图。代码示例:
from rdflib import ConjunctiveGraph, URIRef import json # 原始的JSON-LD文档 JSONLD_DOCUMENT = { "@context": { "schema": "https://schema.org/", "blog": "https://blog.me/", "ex": "https://example.org/", "rdfs": "http://www.w3.org/2000/01/rdf-schema#" }, "@id": "blog:JSONLD-and-named-graphs", "@type": "schema:blogPost", "rdfs:label": "JSON-LD and Named Graphs", "@graph": [ { "@id": "ex:Robot", "@type": "rdfs:Class" }, { "@id": "ex:Rover", "rdfs:subClassOf": { "@id": "ex:Robot" } } ] } target_graph_uri = URIRef('https://myblog.net/rdf/') extra_graph_uri = URIRef('blog:JSONLD-and-named-graphs') # 正常解析文档 graph = ConjunctiveGraph() serialized_document = json.dumps(JSONLD_DOCUMENT) graph.parse( data=serialized_document, format='json-ld', publicID=target_graph_uri, ) # 获取目标图和多余的图 target_graph = graph.get_context(target_graph_uri) extra_graph = graph.get_context(extra_graph_uri) # 把多余图的三元组全部移到目标图 for triple in extra_graph: target_graph.add(triple) # 删除多余的命名图 del graph.store.contexts[extra_graph_uri]
这样处理完,所有数据就都集中在你指定的命名图里了。
不过要提醒一句:这两种方法都是绕过了JSON-LD的原生语义,所以只在你明确需要这种“强制合并”的场景下使用,别随便用,避免破坏数据原本的语义结构哦。
内容的提问来源于stack exchange,提问作者Anatoly Scherbakov
相关产品推荐
相关产品推荐

