You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Google Cloud Platform上Jupyter Lab内核连接问题求助

GCP Jupyter Lab运行音频转录代码时内核崩溃/连接失败问题

我在Google Cloud Platform(GCP)上做音频转录,已创建项目及存储音频的Bucket。在Jupyter Lab中运行下方代码时,最后一步显示“Kernel status: connecting”,一段时间后内核崩溃或持续处于连接状态,有时还会出现“Server Connection Error”提示:

A connection to the Jupyter server could not be established. JupyterLab will continue trying to reconnect. Check your network connection or Jupyter server configuration.

尝试过重启内核和Jupyter,但问题未解决。

运行的代码

# Libraries
#---------------------------------------------------------#
# Data processing libraries
import json

# Transcription library
import whisper

# Environment libraries
import tqdm
import os

# Google Cloud libraries
from google.cloud import storage
import gcsfs

#=========================================================#
# Constants
ID_PROYECTO = 'example'  # Project ID
NOMBRE_BUCKET = 'example-transcript'  # Bucket Name
CARPETA_AUDIOS = 'audios/'  # Audio Folder
CARPETA_TRANSCRIPCIONES = '01. Transcripciones/'  # Transcriptions Folder
IDIOMA = 'es'  # None for autodetect, 'es' for Spanish

# functions to work with buckets
def download_from_gcs(
       client, 
       bucket_name, 
       remote_path, 
       local_path
       ):
   bucket = client.get_bucket(bucket_name)
   blob = bucket.blob(remote_path)
   blob.download_to_filename(local_path)

def save_to_gcs(
       client, 
       bucket_name, 
       content, 
       remote_path, 
       content_type = None
       ):
   bucket = client.get_bucket(bucket_name)
   blob = bucket.blob(remote_path)
   blob.upload_from_string(content, content_type=content_type)

# function for transcription
def transcription(
       path: str = CARPETA_AUDIOS, 
       project_id: str = ID_PROYECTO,
       bucket_name: str = NOMBRE_BUCKET,
       exit_path: str = CARPETA_TRANSCRIPCIONES,
       extension: list = ['mp3', 'wav', 'WAV'],
       model: str = 'large', 
       extract: str = 'all',
       verbose: bool = True,
       language: str = IDIOMA,
       ):
   model = whisper.load_model(model)
   
   client = storage.Client(project = project_id)
   fs = gcsfs.GCSFileSystem(project = project_id)
   all_files = fs.ls(f"{bucket_name}/{path}")
   
   audios = [f for f in all_files if f.split('.')[-1] in extension]

   for i in tqdm.tqdm(audios, disable = verbose is False):
       local_audio_path = os.path.join('/tmp', os.path.basename(i))
       remote_path = i[len(bucket_name)+1:]
       download_from_gcs(client, bucket_name, remote_path, local_audio_path)
       transcript = model.transcribe(
           local_audio_path, 
           language = language
           )
       
       if extract == 'all':
           json_content = json.dumps(transcript)
           json_remote_path = os.path.join(
               exit_path, 
               remote_path+'.json'
               )
           save_to_gcs(
               client,
               bucket_name, 
               json_content, 
               json_remote_path, 
               content_type = 'application/json'
               )

# Function call
transcription()

解决方法

1. 升级Jupyter实例资源

Whisper的large模型对内存、算力要求极高,默认规格的GCP Jupyter实例(如n1-standard-1)不足以支撑:

  • 切换至更高配实例:推荐选择n1-highmem-4(8vCPU/30GB内存)或带T4 GPU的实例,利用GPU加速模型加载与转录,减少内存占用。
  • 监控资源使用:在GCP控制台查看实例的CPU、内存使用率,确认是否因资源耗尽导致内核崩溃。

2. 优化代码减少资源消耗

  • 降级模型测试:先用base或small模型验证代码逻辑正常,再换回large模型。
  • 增加错误处理与资源清理:修改循环部分,添加异常捕获与临时文件清理,避免单个音频处理失败导致进程崩溃,同时释放磁盘空间:
    for i in tqdm.tqdm(audios, disable = verbose is False):
        try:
            local_audio_path = os.path.join('/tmp', os.path.basename(i))
            remote_path = i[len(bucket_name)+1:]
            download_from_gcs(client, bucket_name, remote_path, local_audio_path)
            transcript = model.transcribe(
                local_audio_path, 
                language = language
                )
            
            if extract == 'all':
                json_content = json.dumps(transcript)
                json_remote_path = os.path.join(
                    exit_path, 
                    remote_path+'.json'
                    )
                save_to_gcs(
                    client,
                    bucket_name, 
                    json_content, 
                    json_remote_path, 
                    content_type = 'application/json'
                    )
            # 清理临时文件
            os.remove(local_audio_path)
        except Exception as e:
            print(f"处理音频{i}失败: {str(e)}")
            continue
    

3. 检查网络与权限配置

  • 确认Jupyter实例能访问GCS Bucket:如果使用私有VPC,需确保防火墙规则允许实例与GCS通信(默认GCP实例可直接访问GCS,无需额外配置)。
  • 验证服务账号权限:运行Jupyter的实例服务账号需拥有storage.objects.get和storage.objects.create权限,可绑定Storage Object Admin角色。

4. 修复Jupyter环境

  • 重新安装内核:在终端执行python -m ipykernel install --user,确保内核与当前Python环境匹配。
  • 更新依赖包:执行pip install --upgrade whisper google-cloud-storage gcsfs ipykernel,解决版本冲突问题。

内容的提问来源于stack exchange,提问作者nils

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.05 16:15:04