You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Mediktor爬虫token过期问题及数据提取保存需求

问题解决方案

一、API调用的ME667错误修复(令牌过期)

ME667错误表示用户身份令牌过期,需在请求失败时自动刷新令牌并重试当前请求,同时添加异常捕获避免循环阻塞。

二、数据保存与精简JSON生成

解析API返回的原始数据,提取核心字段生成精简JSON并保存到本地文件。

改进后的完整代码

import json
import requests

def get_auth_code():
    url = "https://www.mediktor.com/vendor.js"
    response = requests.get(url)
    start_index = response.text.index('APP_API_AUTH_CODE:"', 0) + len('APP_API_AUTH_CODE:"')
    end_index = response.text.index('"', start_index)
    return response.text[start_index:end_index]

def get_auth_token_and_device_id():
    url = "https://euapi01.mediktor.com/backoffice/services/login"
    payload = {
        "useCache": 0,
        "apiVersion": "4.1.1",
        "appVersion": "8.7.0",
        "appId": None,
        "deviceType": "WEB",
        "deviceToken": None,
        "language": "pt_BR",
        "timezoneRaw": 180,
        "authTokenRefreshExpiresIn": None
    }
    headers = {
        'authorization': f'Basic {get_auth_code()}',
        'Content-Type': 'application/json'
    }
    response = requests.post(url, headers=headers, json=payload)
    response_data = response.json()
    return response_data['authToken'], response_data['deviceId']

def get_conclusion_list(auth_token, device_id):
    url = "https://euapi01.mediktor.com/backoffice/services/conclusionList"
    payload = {
        "useCache": 168,
        "apiVersion": "4.1.1",
        "appVersion": "8.7.0",
        "appId": None,
        "deviceType": "WEB",
        "deviceToken": None,
        "language": "pt_BR",
        "timezoneRaw": 180,
        "deviceId": device_id
    }
    headers = {
        'accept': 'application/json, text/plain, */*',
        'authorization': f'Bearer {auth_token}',
        'content-type': 'application/json;charset=UTF-8'
    }
    response = requests.post(url, headers=headers, json=payload)
    return [item['conclusionId'] for item in response.json()['conclusions']]

def get_details(conclusion_id, auth_token, device_id):
    url = "https://euapi01.mediktor.com/backoffice/services/conclusionDetail"
    payload = {
        "useCache": 0,
        "apiVersion": "4.1.1",
        "appVersion": "8.7.0",
        "appId": None,
        "deviceType": "WEB",
        "deviceToken": None,
        "language": "pt_BR",
        "timezoneRaw": 180,
        "deviceId": device_id,
        "conclusionId": conclusion_id,
        "conclusionTemplate": "conclusion_description_body",
        "includeActions": True
    }
    headers = {
        'authorization': f'Bearer {auth_token}',
        'content-type': 'application/json;charset=UTF-8'
    }
    response = requests.post(url, headers=headers, json=payload)
    return response.json()

def main():
    auth_token, device_id = get_auth_token_and_device_id()
    conclusion_list = get_conclusion_list(auth_token, device_id)
    simplified_data = []
    output_file = "mediktor_glossary.json"
    
    for idx, conclusion_id in enumerate(conclusion_list):
        print(f"处理第 {idx+1}/{len(conclusion_list)} 条数据...")
        try:
            detail_data = get_details(conclusion_id, auth_token, device_id)
            
            # 检测令牌过期并刷新
            if "error" in detail_data and detail_data["error"]["code"] == "ME667":
                print("令牌过期,重新获取...")
                auth_token, device_id = get_auth_token_and_device_id()
                detail_data = get_details(conclusion_id, auth_token, device_id)
            
            # 提取所需核心字段(可根据需求调整)
            simplified_item = {
                "id": conclusion_id,
                "name": detail_data.get("conclusion", {}).get("name", ""),
                "description": detail_data.get("conclusion", {}).get("description", ""),
                "category": detail_data.get("conclusion", {}).get("categoryName", "")
            }
            simplified_data.append(simplified_item)
            
        except Exception as e:
            print(f"处理ID {conclusion_id} 时出错: {str(e)}")
            continue
    
    # 保存精简JSON
    with open(output_file, "w", encoding="utf-8") as f:
        json.dump(simplified_data, f, ensure_ascii=False, indent=2)
    print(f"数据已保存至 {output_file}")

if __name__ == "__main__":
    main()

代码改进点说明

  • 令牌自动刷新:检测到ME667错误时自动重新获取令牌,并重试当前请求
  • 循环防阻塞:添加全局异常捕获,单个请求失败时跳过继续处理下一条
  • 数据格式优化:用json模块构造请求体,避免字符串拼接的语法错误
  • 语言一致性:将详情请求的语言改为pt_BR,与列表请求保持一致
  • 精简数据提取:提取核心字段生成轻量化JSON,避免冗余数据
  • 进度可视化:添加处理进度提示,方便跟踪任务状态

三、初始Scrapy+Selenium爬虫问题说明

初始爬虫无法正常工作的原因:

  • 重复请求:start_requests已通过Selenium获取所有术语链接并发起请求,parse方法又再次提取链接发起重复请求
  • CSS选择器错误:parse_info中的espec选择器格式错误,多层class需用.连接(如.mdk-ui-list-item__text.mdc-list-item__text span::text),且未处理get()返回None的情况
  • 逻辑错误:desc变量在循环中被反复覆盖,最终仅保留最后一次循环的结果,且未关联对应疾病名称

内容的提问来源于stack exchange,提问作者jvdlb6

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.25 19:45:54