读取URL下载ZIP内.xls文件创建HSSFWorkbook报数组越界如何解决
问题修复方案
错误根因
你代码里存在两个核心问题直接导致了数组越界报错:
ZipEntry.getSize()在通过HTTP流读取ZIP的场景下无法返回正确的文件大小,多数时候会返回-1,你直接用该值创建字节数组会直接出现长度非法问题ZipInputStream.read(byte[])无法保证一次性将整个XLS文件的内容读满字节数组,单次调用只会返回当前读取到的字节数,你拿到的excelBytes是不完整的,传入POI解析时自然会触发数组越界
修复代码
推荐使用ByteArrayOutputStream完整读取ZIP条目内容,无需依赖不可靠的getSize()返回值,同时改用try-with-resources自动管理流关闭避免资源泄漏:
String finalUrl = "https://server/myZip.zip"; URL url = new URL(finalUrl); // 用try-with-resources自动关闭所有流 try (InputStream inputStream = new BufferedInputStream(url.openStream()); ZipInputStream zis = new ZipInputStream(inputStream)) { ZipEntry file; while ((file = zis.getNextEntry()) != null) { if (file.getName().endsWith(".xls")) { log.info("xls file found"); log.info("file name : {}", file.getName()); // 用ByteArrayOutputStream完整读取xls内容 ByteArrayOutputStream baos = new ByteArrayOutputStream(); byte[] buffer = new byte[4096]; int len; while ((len = zis.read(buffer)) > 0) { baos.write(buffer, 0, len); } ByteArrayInputStream excelInputStream = new ByteArrayInputStream(baos.toByteArray()); HSSFWorkbook wb = new HSSFWorkbook(excelInputStream); HSSFSheet sheet = wb.getSheetAt(8); log.info("sheet : {}", sheet.getSheetName()); // 处理完记得关workbook避免POI内存泄漏 wb.close(); excelInputStream.close(); } else { log.info("skip non-xls file: {}", file.getName()); } zis.closeEntry(); } } catch (Exception e) { log.error("读取ZIP内Excel失败", e); }
额外优化点
- 原代码跳过非XLS文件时的日志输出逻辑不合理,会每个非XLS文件都打印一次"xls file not found",修复后改为打印跳过的文件名
- 增加了
zis.closeEntry()显式关闭当前读取完成的ZIP条目,避免后续读取异常 - 如果后续你需要处理.xlsx格式的文件,只需将
HSSFWorkbook替换为XSSFWorkbook即可
内容的提问来源于stack exchange,提问作者Master Shifu
相关产品推荐
相关产品推荐

