You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何通过Python从指定URL获取details.json中的唯一ID3022753?

问题

目标

我需要从指定URL中提取唯一ID,也就是截图里高亮的3022753,这个ID出现在两个位置。

背景

起始URL来自StocksList.txt文件,链接为:https://marketsmithindia.com/mstool/eval/tcs/evaluation.jsp#/。打开该URL后,目标ID在名为details.json?symbol=~的请求里的两个位置:

  • 常规信息 → 请求URL
  • 请求头 → :path:

StocksList.txt内容:

https://marketsmithindia.com/mstool/eval/tcs/evaluation.jsp#/

尝试过程

我试过requests库的各种响应方法,编写了以下Python代码,但未得到预期结果:

import requests
# import pandas as pd

infile = 'C:/Volume/PythonProjects/WIP_Projects/GetJSONResponse/StocksList.txt'
outfile = 'C:/Volume/PythonProjects/WIP_Projects/GetJSONResponse/Response.txt'
timeout = 10
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/103.0.0.0 Safari/537.36"
}

with open(infile, 'r') as IN, open(outfile, 'w') as OUT:
   for DATA in IN:
       url = DATA.strip()
       r = requests.get(url, headers=headers)
       print(r.headers, file=OUT)
       # print(r.json(), file=OUT)
       # print(r.links, file=OUT)

预期输出

3022753 

解决方案

直接请求页面只能获取静态HTML,目标ID是页面加载后通过异步请求(XHR/Fetch)生成的,以下两种方法可解决问题:

方法1:用Selenium模拟浏览器捕获请求

Selenium会模拟真实浏览器加载页面,可捕获所有网络请求并提取ID:

from selenium import webdriver
from selenium.webdriver.chrome.service import Service
from webdriver_manager.chrome import ChromeDriverManager
import time
import json

infile = 'C:/Volume/PythonProjects/WIP_Projects/GetJSONResponse/StocksList.txt'
outfile = 'C:/Volume/PythonProjects/WIP_Projects/GetJSONResponse/Response.txt'

# 初始化无头浏览器
options = webdriver.ChromeOptions()
options.add_argument("--headless=new")
driver = webdriver.Chrome(service=Service(ChromeDriverManager().install()), options=options)

with open(infile, 'r') as IN, open(outfile, 'w') as OUT:
    for DATA in IN:
        url = DATA.strip()
        driver.get(url)
        time.sleep(3)  # 等待页面异步请求加载完成
        
        # 遍历所有网络请求日志
        for log in driver.get_log('performance'):
            msg = log.get('message')
            if 'details.json?symbol=~' in msg:
                log_data = json.loads(msg)
                request_url = log_data['message']['params']['request']['url']
                # 提取URL中的ID
                uid = request_url.split('symbol=~')[-1].split('&')[0]  # 处理可能的后续参数
                print(uid.strip(), file=OUT)
                break

driver.quit()

方法2:分析接口直接请求

用浏览器开发者工具(F12)抓包,找到details.json请求的完整URL、请求头(如Cookie、Referer),直接构造请求获取ID:

import requests

infile = 'C:/Volume/PythonProjects/WIP_Projects/GetJSONResponse/StocksList.txt'
outfile = 'C:/Volume/PythonProjects/WIP_Projects/GetJSONResponse/Response.txt'
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/103.0.0.0 Safari/537.36",
    # 从浏览器抓包复制Cookie、Referer等必要头信息
}

with open(infile, 'r') as IN, open(outfile, 'w') as OUT:
    for DATA in IN:
        # 先请求页面获取必要的会话信息(如Cookie)
        page_url = DATA.strip()
        session = requests.Session()
        session.get(page_url, headers=headers)
        
        # 替换为抓包得到的真实接口URL
        api_url = "https://marketsmithindia.com/mstool/eval/details.json?symbol=~"
        response = session.get(api_url, headers=headers)
        data = response.json()
        
        # 从响应中提取目标ID(需根据实际响应结构调整字段名)
        uid = data.get('对应的ID字段')
        print(uid.strip(), file=OUT)

内容的提问来源于stack exchange,提问作者erukumk

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.13 12:59:53