GCP Dataproc误删Jupyter Notebook如何恢复?GCS无检查点文件夹
First off, let's cut to the chase: if you don't have checkpoints in GCS and haven't set up other safeguards, full recovery isn't guaranteed—but there are still a few avenues to try depending on your cluster's state.
Here's what you can do step by step:
1. Check the Dataproc Master Node's Local Storage
If your Dataproc cluster is still running (not terminated), this is your best first bet. By default, Jupyter notebooks on Dataproc are stored locally on the master node before syncing to GCS (if you set up that sync).
- SSH into the master node (you can do this directly from the GCP Console's Dataproc cluster details page).
- Navigate to the default Jupyter working directory:
cd /home/jupyter/work/ - Even if you deleted the notebook via the Jupyter UI, the local file might still be there—sometimes the GCS sync lags, or the deletion didn't propagate to the local filesystem immediately.
- If the file is gone, you can try Linux file recovery tools like
testdiskorphotorecon the master node's disk. Note: This only works if the disk sectors where the file was stored haven't been overwritten yet, so act fast!
2. Verify GCS Bucket Versioning
If your target GCS bucket had versioning enabled before the notebook was deleted, you can retrieve the deleted file's historical version:
- Go to the GCS Console, open your bucket.
- Click the "Versioning" tab (if you don't see it, enable it first—though this won't retroactively save versions if it wasn't on before).
- Look for the deleted notebook file; you'll see previous versions listed with timestamps. You can restore any of these to your bucket.
Important: If versioning wasn't enabled prior to deletion, this method won't work.
3. Sift Through Cloud Logging for Code Snippets
If your notebook ran any jobs or executed cells that generated logs, Dataproc sends these to Cloud Logging. While this won't give you the full notebook, you might be able to recover critical code snippets:
- Go to Cloud Logging in the GCP Console.
- Use filters like
resource.type="dataproc_cluster" AND logName:"projects/[YOUR_PROJECT]/logs/jupyter"(replace[YOUR_PROJECT]with your project ID) to narrow down relevant logs. - Scan the logs for cell outputs or code that was executed—you can copy these back into a new notebook.
4. Check Jupyter's Local History Database
Jupyter keeps a local history of executed cells in a SQLite database on the master node. If your cluster is still up:
- SSH into the master node and run:
sqlite3 ~/.jupyter/notebook_history.sqlite - Query the history table with:
SELECT * FROM history; - This will show you code snippets from cells you ran across all notebooks. It's not the full notebook structure, but it's better than nothing.
What If None of These Work?
If your cluster has already been terminated, and you didn't have GCS versioning or checkpoints set up, unfortunately, the notebook is likely unrecoverable. Dataproc doesn't keep automatic backups of user-created notebooks unless you explicitly configure them.
Preventive Steps for the Future
To avoid this mess next time:
- Enable GCS bucket versioning for your notebook storage bucket—this lets you recover deleted or overwritten files easily.
- Configure Jupyter to auto-save checkpoints to GCS: When setting up your Dataproc cluster, specify the
--properties spark:spark.jupyter.gcs.checkpoint.dir=gs://[YOUR_BUCKET]/checkpointsflag to ensure checkpoints are synced to GCS. - Manual backups: Periodically download notebooks to your local machine or push them to a Git repo (like Cloud Source Repositories).
- Use Dataproc Workflows: If your notebook is part of a recurring job, convert it to a workflow template so you have a saved, reusable version.
内容的提问来源于stack exchange,提问作者user147529

