如何将超大文本文件中的数值导入Python数组?
嘿,这个问题我之前处理欧洲格式的数值文件时也踩过坑!核心原因是你的数据用逗号作为小数点分隔符,而NumPy的loadtxt()默认只认点号当小数点,所以直接炸锅了。针对600万条数据这种规模,给你几个高效靠谱的解决方案,按需选就行:
方案1:用NumPy的genfromtxt批量转换(贴合原生用法)
genfromtxt比loadtxt更灵活,我们可以先把数据读成字符串数组,再批量替换逗号为点转成浮点型:
import numpy as np # 先按空格分割读成字符串数组 str_arr = np.genfromtxt('your_file.txt', delimiter=' ', dtype=str) # 批量替换逗号为点,转成float数组 arr = np.array([s.replace(',', '.') for s in str_arr], dtype=np.float64)
这个写法不用纠结数据的具体长度,简洁又好理解。
方案2:直接读取文本批量替换(高效直观,首选)
因为你的数据是单行排列,直接一次性读入全部内容,替换逗号后分割转数组就行——600万条数值大概几十MB,内存完全扛得住:
import numpy as np with open('your_file.txt', 'r', encoding='utf-8') as f: # 读入全部内容,把所有逗号换成点 cleaned_content = f.read().replace(',', '.') # 按空格分割成字符串列表,转成NumPy浮点数组 arr = np.array(cleaned_content.split(), dtype=np.float64)
我自己处理这类文件时首选这个方法,代码最少,逻辑最直观。
方案3:分块处理(内存友好型,应对超大型文件)
如果你的文件大到直接读入内存会卡顿(比如几百GB,不过600万条数值大概率到不了),可以分块读取,避免一次性占满内存:
import numpy as np def process_chunk(chunk_str): # 替换逗号为点,分割后转float数组 return np.array(chunk_str.replace(',', '.').split(), dtype=np.float64) arr_list = [] remaining = '' # 存储上一个块末尾不完整的数值 chunk_size = 1024 * 1024 # 每次读1MB,可根据内存调整 with open('your_file.txt', 'r', encoding='utf-8') as f: while True: chunk = f.read(chunk_size) if not chunk: # 处理最后剩下的内容 if remaining: arr_list.append(process_chunk(remaining)) break # 合并上一次的剩余内容和当前块 full_content = remaining + chunk # 找到最后一个空格,确保只处理完整的数值 last_space_pos = full_content.rfind(' ') if last_space_pos == -1: # 当前块没有空格,全部留到下一次 remaining = full_content else: # 处理到最后一个空格前的完整内容 process_part = full_content[:last_space_pos] remaining = full_content[last_space_pos+1:] arr_list.append(process_chunk(process_part)) # 拼接所有分块数组 arr = np.concatenate(arr_list)
这个方法会避免把单个数值拆成两半处理,稳妥性拉满。
总结一下,优先选方案2,简单高效;如果是特别大的文件再用方案3;想贴合NumPy原生用法就用方案1。
内容的提问来源于stack exchange,提问作者Arun Subraminion
相关产品推荐
相关产品推荐

