PubTator API查询处理失败,求助排查500错误原因
问题
我正尝试按照PubTator官方API文档提供的示例代码提交查询,但运行后最终得到500状态码,提示信息为"We have trouble processing your query"。以下是我运行的代码:
text = ('Humboldtia vahliana Wight. is an endangered species belonging to the ' 'family Leguminosae (subfamily Caesalpinioideae). Though the plant propagates ' 'through seeds, low seed setting, seed infestation, poor viability of seed, ' 'reduction in natural regeneration of seedlings as well as anthropogenic ' 'activities like overexploitation, habitat destruction and fragmented ' 'distribution in localities are the major factors hindering the survival of the ' 'species. Hence, strategies must be devised urgently for the conservation of this ' 'species as these recalcitrant seeds do not contribute significantly to the seed ' 'bank. The present attempt was to understand the seed physiology and biochemistry ' 'during embryogenesis and embryo desiccation. H. vahliana seeds took 120 days ' 'after anthesis to acquire full maturity. Immature seeds had higher moisture ' 'content (87.40%) which gradually reduced during maturity and reached 55.42% ' 'at the time of seed shed a true recalcitrant behavior of the seeds. Freshly ' 'fallen mature seeds showed an optimal germination percentage of 82.32% which ' 'was severely affected by the decrease in seed moisture content and the critical ' 'moisture content was found to be 33.63% in which the percentage of germination ' 'was only 30%. Cell membrane damage of seed was found to cause quick loss of ' 'seed viability. The LC-MS/MS analysis showed insignificant amounts of ribose, ' 'arabinose and trehalose but a significant accumulation of fructose in the mature ' 'embryos rather than glucose and sucrose. Embryo drying significantly reduced the ' 'level of these sugars including the stress related trehalose indicating the lack ' 'of biosynthetic machinery to counter desiccation stress in these recalcitrant ' 'seeds.') import requests r = requests.post( 'https://www.ncbi.nlm.nih.gov/research/pubtator-api/annotations/annotate/submit/All', data=text.encode('utf-8')) session_num = r.text status_code = 404 while status_code == 404: result = requests.get( 'https://www.ncbi.nlm.nih.gov/research/pubtator-api/annotations/annotate/retrieve/' + session_num) status_code = result.status_code print(status_code, result.text)
注:API文档说明服务器在处理完成前会返回404状态码,因此编写了while循环持续检查状态码,且确认严格遵循了示例代码的请求格式,未找到该错误的相关说明。
分析与解决建议
- 请求格式不规范:直接提交文本内容时,需显式指定
Content-Type为text/plain,否则服务器可能无法正确解析请求体。修改POST请求如下:headers = {'Content-Type': 'text/plain'} r = requests.post( 'https://www.ncbi.nlm.nih.gov/research/pubtator-api/annotations/annotate/submit/All', data=text, headers=headers) - 会话ID有效性验证:未检查POST请求是否成功就直接提取会话ID,若提交请求本身失败(如返回非200状态码),会导致后续用错误内容作为会话ID查询,引发500错误。添加验证步骤:
if r.status_code != 200: print(f"提交请求失败: {r.status_code} {r.text}") exit() session_num = r.text.strip() - 文本内容或长度问题:提交的文本过长,或包含特殊字符(如百分比符号、括号)可能触发服务器解析异常。先提交一段简短的测试文本(如
"Humboldtia vahliana is a plant."),确认接口正常响应后再逐步扩展内容。 - 服务器临时故障:500错误多为服务器内部问题,可能是负载过高或临时维护。间隔1-2分钟后重试请求,或等待一段时间再测试。
内容的提问来源于stack exchange,提问作者SLotreck
相关产品推荐
相关产品推荐

