如何通过循环将文件夹中的1D numpy数组堆叠为2D数组?
问题描述
需要读取文件夹内296个.txt文件中的1D数组(每个数组含30个元素),将每个1D数组作为行组成一个2D数组。此前手动逐个读取文件并命名为array1、array2……后用np.vstack可完成任务,但使用循环实现时,尝试了vstack、stack、append、concatenate等方法均未成功,创建空2D数组B逐个赋值也无法正常运行。
原代码
import os # Folder Path path = "C:\\2022-07-25 second result more iteration" # Change the directory os.chdir(path) B = np.empty([296, 30], dtype=float) # iterate through all file for file in os.listdir(): # Check whether file is in text format or not if file.endswith(".txt"): file_path = f"{path}\{file}" # open the files A = open(f"{path}\{file}", "r") A = np.loadtxt(A, delimiter="\t") A = np.asarray(A) A = A[:, 0] # because of some reasons the txt files have two similar columns that I just take one of them A = np.subtract(A, n) # just some math (n is another 1D array same size as A) A = np.divide(A, n) # just some math #A = A.T #print("A = ", A) #B = np.vstack (A) #B = np.stack(A) #B = np.stack([A]) #B = np.concatenate([A]) # B = np.append(B, A, axis=0) # print("B in loop = ", B) B[i] = A print("B= ", B) im=plt.imshow(B, origin='lower', extent=[650, 850, 1, 296], aspect='auto', cmap=cm.seismic, norm=colors.CenteredNorm(), interpolation='None')
其中一个txt文件内容(仅取一列,两列内容一致)
6563.64300386213 6563.64300386213 7627.5296220466 7627.5296220466 8922.36588941225 8922.36588941225 10515.1799846774 10515.1799846774 12493.117633046 12493.117633046 14970.0858566 14970.0858566 18090.6895050039 18090.6895050039 22026.1642649415 22026.1642649415 26946.5638244996 26946.5638244996 32926.5439947365 32926.5439947365 39699.096077692 39699.096077692 46279.3014208503 46279.3014208503 50755.8724722489 50755.8724722489 51032.9873391785 51032.9873391785 46565.0683371707 46565.0683371707 39022.807682766 39022.807682766 30795.2702323368 30795.2702323368 23449.1094708325 23449.1094708325 17518.9650853553 17518.9650853553 12968.5269428127 12968.5269428127 9591.79370475353 9591.79370475353 7215.20723743423 7215.20723743423 5791.11928730885 5791.11928730885 5452.1471216442 5452.1471216442 6480.10226555175 6480.10226555175 8910.98447577286 8910.98447577286 11691.5414159922 11691.5414159922 13052.3996945771 13052.3996945771 12588.0161703823 12588.0161703823 11198.2182945086 11198.2182945086
问题分析与解决方案
核心错误
原代码中未定义并递增索引变量i,导致执行B[i] = A时直接报错。此外还有几个可优化的细节:
- 文件路径拼接用
f"{path}\{file}"在Windows下可能因转义字符出现问题,建议用os.path.join更安全 np.loadtxt可直接接收文件路径,无需手动调用openos.listdir()返回的文件顺序不固定,若需要按特定顺序读取,建议对文件列表排序
修正后的代码
import os import numpy as np import matplotlib.pyplot as plt from matplotlib import cm, colors # 文件夹路径 path = "C:\\2022-07-25 second result more iteration" # 初始化索引 i = 0 # 创建空数组 B = np.empty([296, 30], dtype=float) # 获取所有txt文件并排序(可选,保证读取顺序稳定) txt_files = sorted([f for f in os.listdir(path) if f.endswith(".txt")]) for file in txt_files: file_path = os.path.join(path, file) # 直接读取文件,取第一列 A = np.loadtxt(file_path, delimiter="\t")[:, 0] # 执行数学运算(确保n已提前定义且形状与A一致) A = (A - n) / n # 赋值到B的对应行 B[i] = A # 索引递增 i += 1 print("B= ", B) im = plt.imshow(B, origin='lower', extent=[650, 850, 1, 296], aspect='auto', cmap=cm.seismic, norm=colors.CenteredNorm(), interpolation='None') plt.show()
另一种更灵活的方案(无需提前指定数组大小)
如果不确定文件数量,或者不想硬编码296,可以先把所有数组存入列表,最后再拼接成2D数组:
import os import numpy as np import matplotlib.pyplot as plt from matplotlib import cm, colors path = "C:\\2022-07-25 second result more iteration" data_list = [] txt_files = sorted([f for f in os.listdir(path) if f.endswith(".txt")]) for file in txt_files: file_path = os.path.join(path, file) A = np.loadtxt(file_path, delimiter="\t")[:, 0] A = (A - n) / n data_list.append(A) # 拼接成2D数组 B = np.vstack(data_list) print("B= ", B) print("B的形状:", B.shape) im = plt.imshow(B, origin='lower', extent=[650, 850, 1, len(B)], aspect='auto', cmap=cm.seismic, norm=colors.CenteredNorm(), interpolation='None') plt.show()
内容的提问来源于stack exchange,提问作者Nasrin Asgari
相关产品推荐
相关产品推荐

