读取6.25GB CSV文件时触发Pandas OverflowError问题求助
解决读取大CSV文件时的OverflowError问题
当使用pd.read_csv()读取6.25GB的大文件出现OverflowError: Python int too large to convert to C long,而小文件正常时,可尝试以下几种解决方法:
分块读取文件
避免一次性加载整个文件到内存,通过chunksize参数分批次读取后合并:chunk_size = 10**6 # 可根据内存情况调整每次读取的行数 chunk_list = [] for chunk in pd.read_csv(file, chunksize=chunk_size): chunk_list.append(chunk) Data = pd.concat(chunk_list, ignore_index=True)启用低内存模式
设置low_memory=False,让pandas不按分块推断数据类型,减少内存占用并规避溢出问题:Data = pd.read_csv(file, low_memory=False)手动指定数据类型
如果已知各列的数据类型,提前通过dtype参数指定,避免pandas自动推断时产生内存溢出:# 示例:根据实际列名和类型调整 dtype_config = {'big_integer_col': 'float64', 'text_column': 'string'} Data = pd.read_csv(file, dtype=dtype_config)使用大数据处理工具
换用Dask这类专为大数据设计的库,它会自动分块处理大文件:import dask.dataframe as dd Data = dd.read_csv(file).compute()
内容的提问来源于stack exchange,提问作者Mohammed shadab
相关产品推荐
相关产品推荐

