You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何让Jupyter Notebook粘贴文本自动识别Markdown与代码单元格?

将结构化文本自动转换为Jupyter Notebook单元格

实现思路

Jupyter Notebook的.ipynb文件本质是JSON格式,核心结构包含cells数组,每个单元格由cell_type(markdown或code)和source(内容)组成。我们可以通过Python脚本解析目标文本,生成符合格式的JSON文件,直接在Jupyter中打开使用。

解析规则对应

  • # 开头的行:生成顶级Markdown单元格(章节标题)
  • ## 开头的行:生成子标题Markdown单元格(示例标题)
  • 非标题的非空行:归为代码单元格,直到遇到下一个标题或文本结束

具体实现代码

import json

def convert_text_to_notebook(input_text, output_file):
    # 初始化Notebook基础结构
    notebook_template = {
        "cells": [],
        "metadata": {
            "kernelspec": {
                "display_name": "Python 3",
                "language": "python",
                "name": "python3"
            },
            "language_info": {
                "codemirror_mode": {"name": "ipython", "version": 3},
                "file_extension": ".py",
                "mimetype": "text/x-python",
                "name": "python",
                "nbconvert_exporter": "python",
                "pygments_lexer": "ipython3",
                "version": "3.10"
            }
        },
        "nbformat": 4,
        "nbformat_minor": 5
    }

    lines = input_text.splitlines()
    current_code = []

    for line in lines:
        stripped_line = line.strip()
        # 处理章节标题
        if line.startswith("# "):
            # 先提交未完成的代码块
            if current_code:
                notebook_template["cells"].append({
                    "cell_type": "code",
                    "execution_count": None,
                    "metadata": {},
                    "outputs": [],
                    "source": current_code
                })
                current_code = []
            # 添加章节标题单元格
            notebook_template["cells"].append({
                "cell_type": "markdown",
                "metadata": {},
                "source": [line]
            })
        # 处理示例标题
        elif line.startswith("## "):
            # 提交未完成的代码块
            if current_code:
                notebook_template["cells"].append({
                    "cell_type": "code",
                    "execution_count": None,
                    "metadata": {},
                    "outputs": [],
                    "source": current_code
                })
                current_code = []
            # 添加示例标题单元格
            notebook_template["cells"].append({
                "cell_type": "markdown",
                "metadata": {},
                "source": [line]
            })
        # 处理代码内容
        elif stripped_line != "":
            current_code.append(f"{line}\n")
    
    # 处理最后一段代码
    if current_code:
        notebook_template["cells"].append({
            "cell_type": "code",
            "execution_count": None,
            "metadata": {},
            "outputs": [],
            "source": current_code
        })
    
    # 写入Notebook文件
    with open(output_file, "w", encoding="utf-8") as f:
        json.dump(notebook_template, f, indent=2)

# 使用示例
if __name__ == "__main__":
    # 替换为你需要转换的文本内容,或从文件/剪贴板读取
    target_text = """# Code Examples using os module in Python

## Example 1: Get the current working directory using the `os` module

import os

current_working_directory = os.getcwd()
print(f"Current working directory is: {current_working_directory}")

## Example 2: Change the current working directory using the `os` module

import os

path = "/your/desired/directory"
os.chdir(path)
"""
    convert_text_to_notebook(target_text, "os_module_examples.ipynb")

使用方法

  1. 将需要转换的文本替换到target_text变量中,或者通过open()函数读取本地文本文件内容
  2. 运行脚本,生成对应的.ipynb文件
  3. 直接用Jupyter Notebook/Lab打开生成的文件,就能看到自动划分好的Markdown和代码单元格

内容的提问来源于stack exchange,提问作者James

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.22 17:27:44