You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从Python脚本访问JupyterLab笔记本中的Python变量(如Pandas DataFrame)

本地Python脚本访问JupyterLab变量的实现方案

核心思路

利用Jupyter的客户端-内核架构,通过jupyter_client库直接连接正在运行的Jupyter内核,无需在笔记本中添加任何额外代码即可读取变量。

步骤实现

  1. 安装依赖库
    先安装连接内核所需的工具库:

    pip install jupyter_client pandas
    
  2. 定位内核连接文件
    Jupyter会将每个运行中内核的连接信息保存到特定目录:

    • Linux/macOS:~/.local/share/jupyter/runtime/
    • Windows:%APPDATA%\jupyter\runtime\
      目录下的文件命名格式为kernel-xxxx.json,每个文件对应一个正在运行的笔记本内核。
  3. 编写连接脚本
    以下脚本可直接读取目标内核中的变量(以读取名为df的Pandas DataFrame为例):

    from jupyter_client import BlockingKernelClient
    import json
    import os
    import pandas as pd
    
    # 定位内核连接文件(此处以Linux/macOS路径为例,Windows需替换)
    runtime_dir = os.path.expanduser("~/.local/share/jupyter/runtime/")
    kernel_files = [f for f in os.listdir(runtime_dir) if f.startswith("kernel-") and f.endswith(".json")]
    
    # 若存在多个内核,需匹配目标笔记本对应的内核文件
    # 可通过查看内核文件中的pid字段,对应笔记本内核的进程ID来精准定位
    target_kernel_file = os.path.join(runtime_dir, kernel_files[0])
    
    # 加载内核连接信息
    with open(target_kernel_file, 'r') as f:
        conn_info = json.load(f)
    
    # 建立与内核的连接
    client = BlockingKernelClient()
    client.load_connection_info(conn_info)
    client.start_channels()
    
    # 执行代码获取目标变量
    # 先将变量赋值给临时变量,再序列化输出
    exec_msg_id = client.execute("temp_var = df.head()")
    exec_result = client.get_shell_msg(exec_msg_id)
    
    if exec_result['content']['status'] == 'ok':
        # 序列化变量并读取输出
        output_msg_id = client.execute("print(temp_var.to_json(orient='split'))")
        output_result = client.get_shell_msg(output_msg_id)
        var_json = output_result['content']['text'].strip()
        
        # 转换回Pandas DataFrame
        local_df = pd.read_json(var_json, orient='split')
        print("成功获取到的DataFrame:")
        print(local_df)
    else:
        print("读取变量失败,错误信息:")
        print('\n'.join(exec_result['content']['traceback']))
    
    # 关闭连接
    client.stop_channels()
    

注意事项

  • 确保jupyter_client版本与你的JupyterLab版本兼容,版本不匹配可能导致连接失败
  • 若同时运行多个Jupyter内核,需通过内核文件中的pid字段匹配目标笔记本的内核进程ID,避免连接错误的内核
  • 对于大型数据集,使用JSON序列化可能存在性能瓶颈,可替换为pickle(仅在可信环境使用,避免安全风险)
  • 目标内核必须处于运行状态,且对应的笔记本未被关闭

内容的提问来源于stack exchange,提问作者Bertrand

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 19:43:33