You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何通过脚本解析Jupyter Notebook的输入及输出单元格内容?

程序化提取Jupyter Notebook单元格内容方案

完全可以通过脚本直接解析Jupyter Notebook的JSON结构来提取各类输入/输出单元格内容,无需依赖nbconvert或papermill这类工具——因为.ipynb文件本质就是标准JSON格式文本,直接解析最灵活。

一、提取输入单元格内容

Notebook的所有单元格都存在JSON结构的cells字段下,每个单元格通过cell_type区分类型(code/markdown/raw),输入内容存放在source字段(是字符串列表,需拼接成完整文本)。

Python脚本示例

import json

def extract_input_cells(notebook_path):
    # 读取Notebook文件
    with open(notebook_path, 'r', encoding='utf-8') as f:
        nb = json.load(f)
    
    input_cells = []
    for cell in nb['cells']:
        input_cells.append({
            'type': cell['cell_type'],
            'content': ''.join(cell['source'])  # 将列表形式的内容拼接成完整文本
        })
    return input_cells

# 使用示例
notebook_path = 'your_notebook.ipynb'
input_content = extract_input_cells(notebook_path)
for cell in input_content:
    print(f"【{cell['type']}单元格】")
    print(cell['content'])
    print('-' * 50)

这段代码会遍历所有单元格,不管是代码、Markdown还是原始单元格,都能完整提取输入内容。

二、提取输出单元格内容

只有code类型的单元格才有输出内容,存放在outputs字段下。输出有多种类型(控制台流输出、执行结果、错误栈等),需根据output_type字段区分处理:

Python脚本示例

import json

def extract_output_cells(notebook_path):
    with open(notebook_path, 'r', encoding='utf-8') as f:
        nb = json.load(f)
    
    output_cells = []
    for cell_idx, cell in enumerate(nb['cells']):
        if cell['cell_type'] != 'code':
            continue  # 非代码单元格无输出
        
        cell_outputs = []
        for output in cell['outputs']:
            output_info = {'type': output['output_type']}
            
            # 按输出类型提取内容
            if output['output_type'] == 'stream':
                # 控制台输出(stdout/stderr)
                output_info['content'] = ''.join(output['text'])
                output_info['stream_type'] = output['name']
            elif output['output_type'] in ['execute_result', 'display_data']:
                # 执行结果/展示数据,优先提取文本格式
                output_info['content'] = output.get('data', {}).get('text/plain', '无文本输出')
            elif output['output_type'] == 'error':
                # 错误栈信息
                output_info['content'] = '\n'.join(output['traceback'])
            
            cell_outputs.append(output_info)
        
        if cell_outputs:
            output_cells.append({
                'cell_index': cell_idx,
                'cell_input': ''.join(cell['source']),
                'outputs': cell_outputs
            })
    return output_cells

# 使用示例
output_content = extract_output_cells(notebook_path)
for cell in output_content:
    print(f"第{cell['cell_index']}个代码单元格:")
    print("输入内容:")
    print(cell['cell_input'])
    print("输出内容:")
    for output in cell['outputs']:
        print(f"- {output['type']}: {output['content']}")
    print('-' * 50)

如果需要提取图片、HTML等非文本输出,可以从output['data']中获取对应格式的内容(比如output['data']['image/png']是Base64编码的图片数据,可解码保存)。

内容的提问来源于stack exchange,提问作者maciek

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 23:15:21