如何修剪文件路径段,确保各段写入后字节数不超255?
按UTF-8字节数修剪路径段至255字节限制
问题背景
现有代码通过字符数截断路径段,逻辑如下:
// No path component can be longer than 255 chars String[] pathComponents = splitPath(newPath); for(int i=0;i<pathComponents.length - 1;i++) { if (pathComponents[i].length() > MAX_FILELENGTH) { String shortened = pathComponents[i].substring(0, MAX_FILELENGTH - 1); shortened = shortened.trim(); sb.append(shortened).append(File.separator); } else { sb.append(pathComponents[i]).append(File.separator); } }
但当路径段包含多字节UTF-8字符(如中文、emoji等)时,即使字符数≤255,UTF-8编码后的字节数仍可能超过255,导致文件系统写入失败。已能通过以下代码判断字节数是否超限:
if(pathComponents[i].getBytes(StandardCharsets.UTF_8).length > MAX_FILELENGTH)
但需要精准修剪字符,确保修剪后的字符串UTF-8字节数不超过255,且不破坏Unicode字符完整性。
解决方案
实现一个按UTF-8字节数精准修剪的方法,逐个字符计算字节长度,避免截断多字节字符的中间部分:
1. 定义常量与修剪方法
private static final int MAX_FILELENGTH = 255; private static String trimToUtf8ByteLimit(String input) { if (input == null || input.isEmpty()) { return input; } byte[] utf8Bytes = input.getBytes(StandardCharsets.UTF_8); if (utf8Bytes.length <= MAX_FILELENGTH) { return input; } int byteCount = 0; int charIndex = 0; int length = input.length(); while (charIndex < length) { char c = input.charAt(charIndex); int charByteLength; // 计算当前字符的UTF-8字节数 if (c <= 0x7F) { charByteLength = 1; } else if (c <= 0x7FF) { charByteLength = 2; } else if (Character.isHighSurrogate(c)) { // 处理UTF-16代理对(如emoji),占4字节UTF-8 charByteLength = 4; charIndex++; // 跳过低代理字符 } else { charByteLength = 3; } // 累加后超过限制则停止 if (byteCount + charByteLength > MAX_FILELENGTH) { break; } byteCount += charByteLength; charIndex++; } // 截取安全字符范围,可选trim(根据业务需求调整) String trimmed = input.substring(0, charIndex).trim(); // 极端情况二次验证,确保字节数合规 if (trimmed.getBytes(StandardCharsets.UTF_8).length > MAX_FILELENGTH) { return trimToUtf8ByteLimit(trimmed); } return trimmed; }
2. 修改原循环逻辑
String[] pathComponents = splitPath(newPath); StringBuilder sb = new StringBuilder(); // 处理中间路径段 for(int i=0;i<pathComponents.length - 1;i++) { String shortened = trimToUtf8ByteLimit(pathComponents[i]); sb.append(shortened).append(File.separator); } // 处理最后一个路径段(原循环未覆盖) if (pathComponents.length > 0) { String lastComponent = pathComponents[pathComponents.length - 1]; sb.append(trimToUtf8ByteLimit(lastComponent)); }
注意事项
- 代理对处理:针对UTF-16中的代理对字符(如部分emoji),需跳过两个字符,避免截断导致乱码。
- trim可选:若无需去除首尾空格,可移除
trim()步骤,根据业务需求调整。 - 二次验证:trim操作可能改变字节数,增加二次验证确保最终结果合规。
内容的提问来源于stack exchange,提问作者Paul Taylor
相关产品推荐
相关产品推荐

