如何从Python脚本访问JupyterLab笔记本中的Python变量(如Pandas DataFrame)
本地Python脚本访问JupyterLab变量的实现方案
核心思路
利用Jupyter的客户端-内核架构,通过jupyter_client库直接连接正在运行的Jupyter内核,无需在笔记本中添加任何额外代码即可读取变量。
步骤实现
安装依赖库
先安装连接内核所需的工具库:pip install jupyter_client pandas定位内核连接文件
Jupyter会将每个运行中内核的连接信息保存到特定目录:- Linux/macOS:
~/.local/share/jupyter/runtime/ - Windows:
%APPDATA%\jupyter\runtime\
目录下的文件命名格式为kernel-xxxx.json,每个文件对应一个正在运行的笔记本内核。
- Linux/macOS:
编写连接脚本
以下脚本可直接读取目标内核中的变量(以读取名为df的Pandas DataFrame为例):from jupyter_client import BlockingKernelClient import json import os import pandas as pd # 定位内核连接文件(此处以Linux/macOS路径为例,Windows需替换) runtime_dir = os.path.expanduser("~/.local/share/jupyter/runtime/") kernel_files = [f for f in os.listdir(runtime_dir) if f.startswith("kernel-") and f.endswith(".json")] # 若存在多个内核,需匹配目标笔记本对应的内核文件 # 可通过查看内核文件中的pid字段,对应笔记本内核的进程ID来精准定位 target_kernel_file = os.path.join(runtime_dir, kernel_files[0]) # 加载内核连接信息 with open(target_kernel_file, 'r') as f: conn_info = json.load(f) # 建立与内核的连接 client = BlockingKernelClient() client.load_connection_info(conn_info) client.start_channels() # 执行代码获取目标变量 # 先将变量赋值给临时变量,再序列化输出 exec_msg_id = client.execute("temp_var = df.head()") exec_result = client.get_shell_msg(exec_msg_id) if exec_result['content']['status'] == 'ok': # 序列化变量并读取输出 output_msg_id = client.execute("print(temp_var.to_json(orient='split'))") output_result = client.get_shell_msg(output_msg_id) var_json = output_result['content']['text'].strip() # 转换回Pandas DataFrame local_df = pd.read_json(var_json, orient='split') print("成功获取到的DataFrame:") print(local_df) else: print("读取变量失败,错误信息:") print('\n'.join(exec_result['content']['traceback'])) # 关闭连接 client.stop_channels()
注意事项
- 确保
jupyter_client版本与你的JupyterLab版本兼容,版本不匹配可能导致连接失败 - 若同时运行多个Jupyter内核,需通过内核文件中的
pid字段匹配目标笔记本的内核进程ID,避免连接错误的内核 - 对于大型数据集,使用JSON序列化可能存在性能瓶颈,可替换为
pickle(仅在可信环境使用,避免安全风险) - 目标内核必须处于运行状态,且对应的笔记本未被关闭
内容的提问来源于stack exchange,提问作者Bertrand
相关产品推荐
相关产品推荐

