Vaex处理CSV中字符串格式数组列时触发ValueError: setting an array element with a sequence错误的解决方法咨询
Vaex处理CSV中字符串格式数组列时触发ValueError: setting an array element with a sequence错误的解决方法咨询
我手头有一个来自之前项目的CSV文件,需要用Python写脚本处理里面的振动/电信号数据。这个CSV里的DecompressedValue列存储的是长度为16000的浮点数组,但读进来的时候是长字符串格式,必须转成数组才能进行后续分析。
在Pandas里我这么处理完全正常:
import pandas as pd import json signal_df = pd.read_csv('csv_test.csv', sep=';') # 因为DecompressedValue列被读成了长字符串,所以需要用json.loads把每个值转成数组 signal_df.DecompressedValue = signal_df.DecompressedValue.apply(lambda r: json.loads(r))
但换成Vaex实现相同功能时,虽然代码能正常执行,但后续访问DataFrame就会抛出错误。我的Vaex测试代码如下:
import vaex import json test = vaex.from_csv('vaex_test.csv', sep=';') test['DecompressedValue'] = test['DecompressedValue'].apply(lambda r: json.loads(r)) test.head()
运行后触发的ValueError错误信息如下:
[12/19/24 12:50:48] ERROR error evaluating: DecompressedValue at rows 0-5 [dataframe.py]:4101 multiprocessing.pool.RemoteTraceback: """ Traceback (most recent call last): File "c:\Users\user\AppData\Local\anaconda3\envs\py310env\lib\mu ltiprocessing\pool.py", line 125, in worker result = (True, func(*args, **kwds)) File "c:\Users\user\AppData\Local\anaconda3\envs\py310env\lib\si te-packages\vaex\expression.py", line 1629, in _apply result = np.array(result) ValueError: setting an array element with a sequence. The requested array has an inhomogeneous shape after 1 dimensions. The detected shape was (5,) + inhomogeneous part. """
我已经查阅了一些出现相同错误的问题,但它们大多和NumPy数组的不均匀形状相关,感觉我的问题是Vaex的特性导致的。想请教各位,如何在Vaex中正确实现和Pandas一样的效果,把字符串格式的数组列转换为可用的数组?
备注:内容来源于stack exchange,提问作者J. Maria
相关产品推荐
相关产品推荐

