如何让xarray.DataArray.to_zarr输出的zarr文件可被napari读取?
有大型TIFF数组,希望用xarray保存为zarr格式后在napari中查看,但napari无法读取xarray生成的zarr文件。想知道能否通过设置xarray.DataArray.to_zarr的参数解决该问题,倾向于使用xarray而非dask.array.to_zarr。
最小可复现示例
import os import numpy as np import xarray as xa arr = np.random.randint(0, 2**16-1, size=(100, 400, 400)) coords = {'z': np.arange(0, 100), 'y': np.arange(0, 400), 'x': np.arange(0, 400)} da = xa.DataArray(arr, dims=['z', 'y', 'x'], coords=coords) dir_save = os.getcwd() path_save = os.path.join(dir_save, 'test.zarr') da.to_zarr(path_save)
报错信息
File C:\ProgramData\anaconda3\envs\napari_env\lib\site-packages\napari\layers\image_image_utils.py:94, in guess_multiscale(data=[dask.array<from-zarr, shape=(100, 400, 400), dty...hunksize=(25, 100, 100), chunktype=numpy.ndarray>, dask.array<from-zarr, shape=(400,), dtype=int64, chunksize=(400,), chunktype=numpy.ndarray>, dask.array<from-zarr, shape=(400,), dtype=int64, chunksize=(400,), chunktype=numpy.ndarray>, dask.array<from-zarr, shape=(100,), dtype=int64, chunksize=(100,), chunktype=numpy.ndarray>])
93 if not consistent:
---> 94 raise ValueError(
trans = <napari.utils.translations.TranslationBundle object at 0x000002492857E5E0>
sizes = [16000000, 400, 400, 100]
95 trans._(
96 'Input data should be an array-like object, or a sequence of arrays of decreasing size. Got arrays in incorrect order, sizes: {sizes}',
97 deferred=True,
98 sizes=sizes,
99 )
100 )
102 return True, MultiScaleData(data)ValueError: Input data should be an array-like object, or a sequence of arrays of decreasing size. Got arrays in incorrect order, sizes: [16000000, 400, 400, 100]
使用的包版本
- conda-forge napari 0.5.5 hd8ed1ab_0
- conda-forge napari-base 0.5.5 pyh9208f05_0
- conda-forge napari-console 0.1.3 pyh73487a3_0
- conda-forge napari-plugin-engine 0.2.0 pyha07c04f_3
- conda-forge napari-plugin-manager 0.1.4 pyha07c04f_0
- conda-forge napari-svg 0.2.1 pyha07c04f_0
- conda-forge xarray 2025.1.1 pyhd8ed1ab_0
解决方案
问题根源
xarray默认会将DataArray的坐标(z、y、x)作为独立数组写入zarr目录中。napari读取zarr目录时,会把目录内所有数组视为多尺度图像的不同层级,误将主图像数组和三个坐标数组当成多尺度序列,因尺寸顺序不符合要求触发报错。
修改代码
调用to_zarr时添加exclude参数,排除坐标变量,只保存主图像数组:
import os import numpy as np import xarray as xa arr = np.random.randint(0, 2**16-1, size=(100, 400, 400)) coords = {'z': np.arange(0, 100), 'y': np.arange(0, 400), 'x': np.arange(0, 400)} da = xa.DataArray(arr, dims=['z', 'y', 'x'], coords=coords) dir_save = os.getcwd() path_save = os.path.join(dir_save, 'test.zarr') # 排除坐标变量,仅保存主数据 da.to_zarr(path_save, exclude=['z', 'y', 'x'], mode='w')
参数说明
exclude=['z', 'y', 'x']:指定不写入这三个坐标变量,确保zarr目录中只有主图像数组。mode='w':以写入模式操作,覆盖已有zarr文件,避免旧数据残留干扰读取。
这样生成的zarr文件就能被napari正常识别并打开。
内容的提问来源于stack exchange,提问作者user8188435

