如何无需存储到磁盘,直接将(Geo)TIFF文件读取为numpy数组?
I often download (Geo)TIFF files, save them to temporary disk space, then use rasterio to read the data into a numpy.ndarray for analysis. For example, here's my code for NAIP imagery URLs:
import os from requests import get from rasterio import open as rasopen req = get(url, verify=False, stream=True) if req.status_code != 200: raise ValueError('Bad response from NAIP API request.') temp = os.path.join(os.getcwd(), 'temp', 'tile.tif') with open(temp, 'wb') as f: f.write(req.content)
I want to skip saving the file to disk entirely and directly convert the downloaded (Geo)TIFF into a numpy array. How can I do this?
Great question—you can cut out the disk I/O entirely by using Python's io.BytesIO module to handle the TIFF content in memory. This creates a file-like object that rasterio can read directly, no temporary files required.
Here's a streamlined version of your code that does this:
import io import requests from rasterio import open as rasopen # Fetch the TIFF content from the API req = requests.get(url, verify=False) if req.status_code != 200: raise ValueError('Bad response from NAIP API request.') # Wrap the raw bytes in an in-memory file buffer with rasopen(io.BytesIO(req.content)) as src: # Read all bands into a numpy array tiff_array = src.read() # You can still access geospatial metadata if needed crs = src.crs transform = src.transform
What's happening here:
io.BytesIOtakes the raw byte content from your API response and turns it into a file-like interface. Rasterio’sopen()function accepts this just like it would a regular file path on disk.- The
src.read()call pulls the data directly from the in-memory buffer into anumpy.ndarray(shaped as(number_of_bands, height, width)), ready for your analysis. - All the geospatial metadata you’d get from a disk-based TIFF (like CRS, coordinate transform, band info) is still accessible via the
srcobject.
If you’re dealing with large TIFFs and want to keep using streaming to avoid loading the entire file into memory at once, you can stream chunks into the BytesIO buffer incrementally:
import io import requests from rasterio import open as rasopen req = requests.get(url, verify=False, stream=True) if req.status_code != 200: raise ValueError('Bad response from NAIP API request.') # Stream content into the in-memory buffer in chunks with io.BytesIO() as buffer: for chunk in req.iter_content(chunk_size=8192): buffer.write(chunk) buffer.seek(0) # Reset the buffer to the start so rasterio can read it with rasopen(buffer) as src: tiff_array = src.read()
This approach keeps everything in memory, making your workflow faster and avoiding clutter from temporary files.
内容的提问来源于stack exchange,提问作者dgketchum

