使用readlines()加载2.6GB文件时遇OSError: [Errno 22]无效参数求助
OSError: [Errno 22] Invalid Argument When Loading 2.6GB Extensionless Files with readlines()
问题详情
- 调用Python的
readlines()加载约2.6GB的无扩展名文件(adb2d_x_y、dens2d_x_y)时触发OSError: [Errno 22] Invalid argument错误 - 相同文件、代码逻辑在其他目录下可正常执行
- 文件来自Linux小众程序,可通过文本编辑器打开
- 在文件同目录的Jupyter Notebook中运行代码,已用
os.getcwd()和os.path.join()拼接路径,问题未解决
错误栈
--------------------------------------------------------------------------- OSError Traceback (most recent call last) ~\AppData\Local\Temp\ipykernel_19332\1902527957.py in <module> ----> 1 data=wf_loader('adb', state=2) 2 #wf_loader('adb', state=1) 3 #wf_loader('diab', state=1) 4 #wf_loader('diab', state=2) ~\AppData\Local\Temp\ipykernel_19332\168813414.py in wf_loader(representation, state, save) 22 with open(file) as di_data: 23 print(di_data) ---> 24 lines = di_data.readlines() 25 26 xlims, ylims, x_gridpoints, y_gridpoints = get_data_shape_fromlog() OSError: [Errno 22] Invalid argument
相关代码片段
def wf_loader(representation='adb', state=2, save=True): ''' Returns a numpy array for the wavefunction density progressing in time representation: 'adb' opens the adiabatic wf, 'diab opens the diabatic wf', default adb state: which state's wavefunction to look at 1 or 2, default 2 save: saves the numpy array as a binary file for easier loading later ''' cwd = os.getcwd() if representation == 'adb': file = os.path.join(cwd, 'adb2d_x_y') elif representation == 'diab': file = os.path.join(cwd, 'dens2d_x_y') else: print('Unkown representation') return #Loads file into python with open(file) as di_data: print(di_data) lines = di_data.readlines()
针对性解决方案
1. 替换readlines()为逐行读取
readlines()会一次性将整个2.6GB文件加载到内存,极易触发系统或内存限制。改用迭代器逐行读取,大幅降低内存占用:
with open(file) as di_data: for line in di_data: # 在这里添加单条行的处理逻辑 # 例如:parse_line(line)
2. 显式指定文件打开模式与编码
无扩展名文件可能导致系统默认模式/编码识别异常,尝试显式指定:
# 尝试兼容更广的latin-1编码(如果是文本文件) with open(file, 'r', encoding='latin-1') as di_data: for line in di_data: # 处理逻辑 pass # 如果文件是二进制格式,用rb模式读取后解码 with open(file, 'rb') as di_data: for line in di_data: decoded_line = line.decode('utf-8') # 替换为文件实际编码 # 处理逻辑 pass
3. 检查文件系统与目录权限
当前目录可能存在特殊限制(如Windows加密目录、网络共享盘),将文件复制到本地普通目录(如D:\temp)后重新测试。
4. 使用内存映射处理大文件
通过mmap模块将文件映射到内存,避免一次性加载全量数据:
import mmap with open(file, 'r') as f: # 创建内存映射 with mmap.mmap(f.fileno(), length=0, access=mmap.ACCESS_READ) as mm: # 逐行读取映射内容 for line in iter(mm.readline, b''): decoded_line = line.decode('utf-8') # 处理逻辑 pass
5. 验证文件完整性
对比当前目录文件与正常目录文件的哈希值,确认文件未损坏:
import hashlib def get_file_hash(file_path): sha256 = hashlib.sha256() with open(file_path, 'rb') as f: while chunk := f.read(4096): sha256.update(chunk) return sha256.hexdigest() # 计算当前文件哈希 print(get_file_hash(file)) # 与正常目录下的文件哈希值对比
内容的提问来源于stack exchange,提问作者Nemo
相关产品推荐
相关产品推荐

