Python访问Google Cloud Storage存在文件时触发FileNotFoundError求助
Hey Sara, let's break down why you're hitting that FileNotFoundError and how to fix it quickly.
The Root Cause
Python's built-in os module tools (like listdir, isfile, and join) only work with local file system paths or directories mounted to your machine. They don't natively recognize the gs:// protocol used by Google Cloud Storage—so when your code tries to look up gs://agriculture-bucket-gl/Data sets/, it treats it like a local folder that doesn't exist, hence the error.
Here are two solid solutions to get you accessing your bucket content:
Solution 1: Use the Official GCS Python Client Library (Recommended)
This is the most reliable way to interact with GCS directly from Python, no workarounds needed.
First, install the library (if you haven't already)
Run this in a Jupyter cell to add the Google Cloud Storage client:
!pip install google-cloud-storage
Rewrite your get_files function for GCS
Replace your os-based code with this, which uses the GCS API to list files:
from google.cloud import storage def get_files(bucket_name, folder_prefix=""): # Initialize the GCS client (authenticates automatically if your cluster is linked to GCP) client = storage.Client() bucket = client.get_bucket(bucket_name) # List all "blobs" (GCS terms for files/folders) in the target folder blobs = bucket.list_blobs(prefix=folder_prefix) for blob in blobs: # Skip placeholder "folder" entries (GCS uses prefixes, not actual folders) if not blob.name.endswith('/'): print("File path:", blob.name) # Call it with your bucket and folder path get_files("agriculture-bucket-gl", "Data sets/")
Bonus: Read/write files directly from GCS
Need to pull a file into your notebook? Use this:
def read_gcs_file(bucket_name, file_path): client = storage.Client() bucket = client.get_bucket(bucket_name) blob = bucket.blob(file_path) # Read as plain text (use download_as_bytes() for binary files) return blob.download_as_text() # Example: Read a CSV from your bucket csv_content = read_gcs_file("agriculture-bucket-gl", "Data sets/sample_data.csv")
Uploading a local file to GCS is just as easy:
def upload_to_gcs(local_file_path, bucket_name, gcs_destination_path): client = storage.Client() bucket = client.get_bucket(bucket_name) blob = bucket.blob(gcs_destination_path) blob.upload_from_filename(local_file_path) # Example: Upload a local CSV to your bucket upload_to_gcs("./my_local_file.csv", "agriculture-bucket-gl", "Data sets/uploaded_file.csv")
Solution 2: Mount the GCS Bucket as a Local Folder with gcsfuse
If you want to keep using your original os-based code, you can mount the GCS bucket to your cluster's local file system using gcsfuse. This makes the bucket act like a regular folder on your machine.
Step 1: Install gcsfuse (skip if already installed)
Run these commands in a Jupyter cell or cluster terminal:
!echo "deb http://packages.cloud.google.com/apt gcsfuse-$(lsb_release -c -s) main" | sudo tee /etc/apt/sources.list.d/gcsfuse.list !curl https://packages.cloud.google.com/apt/doc/apt-key.gpg | sudo apt-key add - !sudo apt-get update !sudo apt-get install gcsfuse -y
Step 2: Mount the bucket
Create a local directory and mount your GCS bucket to it:
!mkdir -p /mnt/my_gcs_bucket !gcsfuse agriculture-bucket-gl /mnt/my_gcs_bucket
Step 3: Use your original code with the mounted path
Now you can reference the bucket via the local mount path instead of gs://:
import os from os import listdir from os.path import isfile, join # Point to the mounted folder instead of the gs:// path localFolder = "/mnt/my_gcs_bucket/Data sets/" def get_files(bucketName): files = [f for f in listdir(localFolder) if isfile(join(localFolder, f))] for file in files: print("file path:", file) get_files("agriculture-bucket-gl")
Note: If you're using Vertex AI Workbench or a managed Dataproc cluster, gcsfuse might already be pre-installed, so you can jump straight to mounting.
内容的提问来源于stack exchange,提问作者SaraAwad

