Python中SPARQL查询变量拼接报错,寻求解决方法
问题描述
编写Python函数调用SPARQL端点,意图根据用户输入修改查询中mesh:"+d+" meshv:treeNumber ?treeNum .部分,但运行时出现SPARQL语法错误,报错信息如下:
JSONDecodeError: [Errno Expecting value] Encountered " <STRING_LITERAL2> ""+d+" "" at line 16, column 14.
Was expecting one of:...
<PNAME_NS> ...
<PNAME_LN> ...... ...
"a" ...
"distinct" ...
"multi" ...
"shortest" ...
"(" ...
"!" ...
"^" ...
: 0
原函数代码:
def get_disease(d): query = ''' PREFIX rdf: <http://www.w3.org/1999/02/22-rdf-syntax-ns#> PREFIX rdfs: <http://www.w3.org/2000/01/rdf-schema#> PREFIX xsd: <http://www.w3.org/2001/XMLSchema#> PREFIX owl: <http://www.w3.org/2002/07/owl#> PREFIX meshv: <http://id.nlm.nih.gov/mesh/vocab#> PREFIX mesh: <http://id.nlm.nih.gov/mesh/> PREFIX mesh2022: <http://id.nlm.nih.gov/mesh/2022/> SELECT DISTINCT ?descriptor ?label FROM <http://id.nlm.nih.gov/mesh> WHERE { mesh:"+d+" meshv:treeNumber ?treeNum . ?childTreeNum meshv:parentTreeNumber+ ?treeNum . ?descriptor meshv:treeNumber ?childTreeNum . ?descriptor rdfs:label ?label . } ORDER BY ?label ''' #set the url url = 'https://id.nlm.nih.gov/mesh/sparql' #set the header headers = {'Content-Type': 'application/sparql-results+json'} #set the parameters params = {'query' : query, 'limit' : 1000, 'inference' : 'true', 'format' : 'JSON'} #send the request response = requests.get(url, headers = headers, params = params) jsonResponse = response.json()['results'] get_disease('D020521')
问题原因
Python的三引号字符串('''/""")不会解析其中的"+d+"表达式,导致最终发送给SPARQL端点的查询语句里直接保留了mesh:"+d+"这段无效语法,SPARQL引擎无法识别,因此抛出语法错误。
解决方案
有两种可靠的方式实现参数替换:
方式1:使用Python的f-string格式化
将查询字符串改为f-string,直接在对应位置插入变量d:
def get_disease(d): query = f''' PREFIX rdf: <http://www.w3.org/1999/02/22-rdf-syntax-ns#> PREFIX rdfs: <http://www.w3.org/2000/01/rdf-schema#> PREFIX xsd: <http://www.w3.org/2001/XMLSchema#> PREFIX owl: <http://www.w3.org/2002/07/owl#> PREFIX meshv: <http://id.nlm.nih.gov/mesh/vocab#> PREFIX mesh: <http://id.nlm.nih.gov/mesh/> PREFIX mesh2022: <http://id.nlm.nih.gov/mesh/2022/> SELECT DISTINCT ?descriptor ?label FROM <http://id.nlm.nih.gov/mesh> WHERE {{ mesh:{d} meshv:treeNumber ?treeNum . ?childTreeNum meshv:parentTreeNumber+ ?treeNum . ?descriptor meshv:treeNumber ?childTreeNum . ?descriptor rdfs:label ?label . }} ORDER BY ?label ''' url = 'https://id.nlm.nih.gov/mesh/sparql' headers = {'Content-Type': 'application/sparql-results+json'} params = {'query' : query, 'limit' : 1000, 'inference' : 'true', 'format' : 'JSON'} response = requests.get(url, headers = headers, params = params) jsonResponse = response.json()['results'] get_disease('D020521')
注意:f-string中如果需要保留大括号{},需要写成双大括号{{}},否则会被解析为占位符。
方式2:使用字符串format()方法
如果Python版本不支持f-string(Python 3.6以下),可以用str.format():
def get_disease(d): query = ''' PREFIX rdf: <http://www.w3.org/1999/02/22-rdf-syntax-ns#> PREFIX rdfs: <http://www.w3.org/2000/01/rdf-schema#> PREFIX xsd: <http://www.w3.org/2001/XMLSchema#> PREFIX owl: <http://www.w3.org/2002/07/owl#> PREFIX meshv: <http://id.nlm.nih.gov/mesh/vocab#> PREFIX mesh: <http://id.nlm.nih.gov/mesh/> PREFIX mesh2022: <http://id.nlm.nih.gov/mesh/2022/> SELECT DISTINCT ?descriptor ?label FROM <http://id.nlm.nih.gov/mesh> WHERE {{ mesh:{0} meshv:treeNumber ?treeNum . ?childTreeNum meshv:parentTreeNumber+ ?treeNum . ?descriptor meshv:treeNumber ?childTreeNum . ?descriptor rdfs:label ?label . }} ORDER BY ?label '''.format(d) url = 'https://id.nlm.nih.gov/mesh/sparql' headers = {'Content-Type': 'application/sparql-results+json'} params = {'query' : query, 'limit' : 1000, 'inference' : 'true', 'format' : 'JSON'} response = requests.get(url, headers = headers, params = params) jsonResponse = response.json()['results'] get_disease('D020521')
额外建议
如果参数是用户输入的非可信内容,直接拼接字符串存在注入风险,建议使用SPARQL参数绑定机制(如果端点支持),不过对于这个特定的Mesh端点,上述两种方式已经能满足需求且安全,因为输入是Mesh标识符(如D020521),格式固定。
内容的提问来源于stack exchange,提问作者rshar

