You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用readlines()加载2.6GB文件时遇OSError: [Errno 22]无效参数求助

OSError: [Errno 22] Invalid Argument When Loading 2.6GB Extensionless Files with readlines()

问题详情

  • 调用Python的readlines()加载约2.6GB的无扩展名文件(adb2d_x_y、dens2d_x_y)时触发OSError: [Errno 22] Invalid argument错误
  • 相同文件、代码逻辑在其他目录下可正常执行
  • 文件来自Linux小众程序,可通过文本编辑器打开
  • 在文件同目录的Jupyter Notebook中运行代码,已用os.getcwd()和os.path.join()拼接路径,问题未解决

错误栈

---------------------------------------------------------------------------
OSError                                   Traceback (most recent call last)
~\AppData\Local\Temp\ipykernel_19332\1902527957.py in <module>
----> 1 data=wf_loader('adb', state=2)
      2 #wf_loader('adb', state=1)
      3 #wf_loader('diab', state=1)
      4 #wf_loader('diab', state=2)

~\AppData\Local\Temp\ipykernel_19332\168813414.py in wf_loader(representation, state, save)
     22     with open(file) as di_data:
     23         print(di_data)
---&gt; 24         lines = di_data.readlines()
     25 
     26     xlims, ylims, x_gridpoints, y_gridpoints = get_data_shape_fromlog()

OSError: [Errno 22] Invalid argument

相关代码片段

def wf_loader(representation='adb', state=2, save=True):
    '''
    Returns a numpy array for the wavefunction density progressing in time
    
    representation: 'adb' opens the adiabatic wf, 'diab opens the diabatic wf', default adb
    
    state: which state's wavefunction to look at 1 or 2, default 2
    
    save: saves the numpy array as a binary file for easier loading later
    '''

    cwd = os.getcwd()
    
    if representation == 'adb':
        file = os.path.join(cwd, 'adb2d_x_y')
    elif representation == 'diab':
        file = os.path.join(cwd, 'dens2d_x_y')
    else:
        print('Unkown representation')
        return
        
    #Loads file into python
    with open(file) as di_data:
        print(di_data)
        lines = di_data.readlines()

针对性解决方案

1. 替换readlines()为逐行读取

readlines()会一次性将整个2.6GB文件加载到内存,极易触发系统或内存限制。改用迭代器逐行读取,大幅降低内存占用:

with open(file) as di_data:
    for line in di_data:
        # 在这里添加单条行的处理逻辑
        # 例如:parse_line(line)

2. 显式指定文件打开模式与编码

无扩展名文件可能导致系统默认模式/编码识别异常,尝试显式指定:

# 尝试兼容更广的latin-1编码(如果是文本文件)
with open(file, 'r', encoding='latin-1') as di_data:
    for line in di_data:
        # 处理逻辑
        pass

# 如果文件是二进制格式,用rb模式读取后解码
with open(file, 'rb') as di_data:
    for line in di_data:
        decoded_line = line.decode('utf-8') # 替换为文件实际编码
        # 处理逻辑
        pass

3. 检查文件系统与目录权限

当前目录可能存在特殊限制(如Windows加密目录、网络共享盘),将文件复制到本地普通目录(如D:\temp)后重新测试。

4. 使用内存映射处理大文件

通过mmap模块将文件映射到内存,避免一次性加载全量数据:

import mmap

with open(file, 'r') as f:
    # 创建内存映射
    with mmap.mmap(f.fileno(), length=0, access=mmap.ACCESS_READ) as mm:
        # 逐行读取映射内容
        for line in iter(mm.readline, b''):
            decoded_line = line.decode('utf-8')
            # 处理逻辑
            pass

5. 验证文件完整性

对比当前目录文件与正常目录文件的哈希值,确认文件未损坏:

import hashlib

def get_file_hash(file_path):
    sha256 = hashlib.sha256()
    with open(file_path, 'rb') as f:
        while chunk := f.read(4096):
            sha256.update(chunk)
    return sha256.hexdigest()

# 计算当前文件哈希
print(get_file_hash(file))
# 与正常目录下的文件哈希值对比

内容的提问来源于stack exchange,提问作者Nemo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.27 05:12:36