如何将Vaex DataFrame转换为NumPy数组?遇NotImplementedError求助
解决Vaex DataFrame转NumPy数组的NotImplementedError问题
你遇到的NotImplementedError是因为Vaex的列对象(Expression)不支持Pandas风格的.values属性。Vaex为了实现大数据的延迟计算,采用了和Pandas不同的设计逻辑,需要用它专属的方法来提取NumPy数组。
修正方案
有两种可行的方式解决这个问题:
1. 对单个列使用to_numpy()方法
将原来的.values替换为Vaex提供的to_numpy()方法,这是官方推荐的获取单列NumPy数组的方式:
# 替换原来的列提取代码 x = df['x'].to_numpy() y = df['y'].to_numpy() z = df['z'].to_numpy()
2. 直接将整个DataFrame转为NumPy数组(更高效)
如果最终需要的是合并后的(n,3)数组,直接调用DataFrame的to_numpy()方法可以一步到位,避免逐个列转换的冗余操作,尤其适合大数据场景:
points = df.to_numpy()
修改后的完整函数
替换错误部分后的read_pts_file函数如下:
# Function to read pts file using Vaex def read_pts_file(filename): # Create a temporary file with tempfile.NamedTemporaryFile(delete=False, mode='w', newline='') as temp_file: # Read the original .pts file and write to the temporary file while skipping the first row with open(filename, 'r') as original_file: original_file.readline() # Skip the first line for line in original_file: temp_file.write(line) # Write the rest of the lines temp_filename = temp_file.name # Save the name of the temporary file print("Input file path:", filename) print("Temporary file path:", temp_filename) # Use Vaex to read from the temporary file df = vaex.from_csv(temp_filename, sep=' ', convert=True, header=None) # Only take the first 3 columns of the data df = df[[0, 1, 2]] # Rename columns for clarity (optional) df.columns = ['x', 'y', 'z'] # 直接将整个DataFrame转为NumPy数组,一步得到(n,3)的结果 points = df.to_numpy() # Clean up the temporary file os.remove(temp_filename) # Return points and color data return points
补充说明
Vaex的核心设计是延迟计算,所有操作不会立即执行,直到你显式请求数据(比如调用to_numpy()、head()等方法)。这和Pandas的即时计算逻辑不同,所以不能直接照搬Pandas的属性和方法。
内容的提问来源于stack exchange,提问作者Conan Moon
相关产品推荐
相关产品推荐

