Mediktor爬虫token过期问题及数据提取保存需求
问题解决方案
一、API调用的ME667错误修复(令牌过期)
ME667错误表示用户身份令牌过期,需在请求失败时自动刷新令牌并重试当前请求,同时添加异常捕获避免循环阻塞。
二、数据保存与精简JSON生成
解析API返回的原始数据,提取核心字段生成精简JSON并保存到本地文件。
改进后的完整代码
import json import requests def get_auth_code(): url = "https://www.mediktor.com/vendor.js" response = requests.get(url) start_index = response.text.index('APP_API_AUTH_CODE:"', 0) + len('APP_API_AUTH_CODE:"') end_index = response.text.index('"', start_index) return response.text[start_index:end_index] def get_auth_token_and_device_id(): url = "https://euapi01.mediktor.com/backoffice/services/login" payload = { "useCache": 0, "apiVersion": "4.1.1", "appVersion": "8.7.0", "appId": None, "deviceType": "WEB", "deviceToken": None, "language": "pt_BR", "timezoneRaw": 180, "authTokenRefreshExpiresIn": None } headers = { 'authorization': f'Basic {get_auth_code()}', 'Content-Type': 'application/json' } response = requests.post(url, headers=headers, json=payload) response_data = response.json() return response_data['authToken'], response_data['deviceId'] def get_conclusion_list(auth_token, device_id): url = "https://euapi01.mediktor.com/backoffice/services/conclusionList" payload = { "useCache": 168, "apiVersion": "4.1.1", "appVersion": "8.7.0", "appId": None, "deviceType": "WEB", "deviceToken": None, "language": "pt_BR", "timezoneRaw": 180, "deviceId": device_id } headers = { 'accept': 'application/json, text/plain, */*', 'authorization': f'Bearer {auth_token}', 'content-type': 'application/json;charset=UTF-8' } response = requests.post(url, headers=headers, json=payload) return [item['conclusionId'] for item in response.json()['conclusions']] def get_details(conclusion_id, auth_token, device_id): url = "https://euapi01.mediktor.com/backoffice/services/conclusionDetail" payload = { "useCache": 0, "apiVersion": "4.1.1", "appVersion": "8.7.0", "appId": None, "deviceType": "WEB", "deviceToken": None, "language": "pt_BR", "timezoneRaw": 180, "deviceId": device_id, "conclusionId": conclusion_id, "conclusionTemplate": "conclusion_description_body", "includeActions": True } headers = { 'authorization': f'Bearer {auth_token}', 'content-type': 'application/json;charset=UTF-8' } response = requests.post(url, headers=headers, json=payload) return response.json() def main(): auth_token, device_id = get_auth_token_and_device_id() conclusion_list = get_conclusion_list(auth_token, device_id) simplified_data = [] output_file = "mediktor_glossary.json" for idx, conclusion_id in enumerate(conclusion_list): print(f"处理第 {idx+1}/{len(conclusion_list)} 条数据...") try: detail_data = get_details(conclusion_id, auth_token, device_id) # 检测令牌过期并刷新 if "error" in detail_data and detail_data["error"]["code"] == "ME667": print("令牌过期,重新获取...") auth_token, device_id = get_auth_token_and_device_id() detail_data = get_details(conclusion_id, auth_token, device_id) # 提取所需核心字段(可根据需求调整) simplified_item = { "id": conclusion_id, "name": detail_data.get("conclusion", {}).get("name", ""), "description": detail_data.get("conclusion", {}).get("description", ""), "category": detail_data.get("conclusion", {}).get("categoryName", "") } simplified_data.append(simplified_item) except Exception as e: print(f"处理ID {conclusion_id} 时出错: {str(e)}") continue # 保存精简JSON with open(output_file, "w", encoding="utf-8") as f: json.dump(simplified_data, f, ensure_ascii=False, indent=2) print(f"数据已保存至 {output_file}") if __name__ == "__main__": main()
代码改进点说明
- 令牌自动刷新:检测到ME667错误时自动重新获取令牌,并重试当前请求
- 循环防阻塞:添加全局异常捕获,单个请求失败时跳过继续处理下一条
- 数据格式优化:用
json模块构造请求体,避免字符串拼接的语法错误 - 语言一致性:将详情请求的语言改为
pt_BR,与列表请求保持一致 - 精简数据提取:提取核心字段生成轻量化JSON,避免冗余数据
- 进度可视化:添加处理进度提示,方便跟踪任务状态
三、初始Scrapy+Selenium爬虫问题说明
初始爬虫无法正常工作的原因:
- 重复请求:
start_requests已通过Selenium获取所有术语链接并发起请求,parse方法又再次提取链接发起重复请求 - CSS选择器错误:
parse_info中的espec选择器格式错误,多层class需用.连接(如.mdk-ui-list-item__text.mdc-list-item__text span::text),且未处理get()返回None的情况 - 逻辑错误:
desc变量在循环中被反复覆盖,最终仅保留最后一次循环的结果,且未关联对应疾病名称
内容的提问来源于stack exchange,提问作者jvdlb6
相关产品推荐
相关产品推荐

