Java读取S3下载文件时特殊字符丢失问题求助
S3文件特殊字符(ä、ü)读取后变为空格的解决方法
S3存储的文件包含ä、ü这类特殊字符,Java下载后读取/写入时这些字符会被替换为空空格。尝试过多种UTF-8格式读写方式(BufferedReader、BufferedWriter、FileReader、FileWriter等)均无效,且已确认S3下载环节无问题,问题出在Java读取后的处理流程中。
当前代码片段
Java下载代码
String uploadPath = "/u01/infor/sce/scp11/logs/temp.sql"; AmazonS3 s3Client = AmazonS3ClientBuilder.defaultClient(); S3Object obj = s3Client.getObject(System.getenv("BUCKET"), "/temp.sql"); S3ObjectInputStream s3ObjectInputStream = obj.response().asInputStream(); // 下载到本地文件 String localFilePath = uploadPath; try (OutputStream outputStream = new FileOutputStream(localFilePath)) { byte[] buffer = new byte[1024]; int bytesRead; while ((bytesRead = s3ObjectInputStream.read(buffer)) != -1) { outputStream.write(buffer, 0, bytesRead); } System.out.println("File downloaded successfully."); } catch (Exception e) { log.info("Error while downloading sql file from s3" + e.getMessage(), e); }
可正常处理特殊字符的Python代码
import io import json import boto3 import botocore import datetime import base64 import os import random import string import time s3 = boto3.resource('s3') s3_object = boto3.resource('s3').Object('hg-us-east-1', 'logs/delete-entries.sql') data = io.BytesIO() s3_object.download_fileobj(data) s = None b = s3_object.get()['Body'] s = b.read().decode('utf-8') print(s) # 写入时明确指定UTF-8编码,避免依赖系统默认 f = open("c:\\Users\\tim\\test.sql", "w", encoding='utf-8') f.write(s) f.close()
Java读取解析代码
import java.io.BufferedReader; import java.io.FileInputStream; import java.io.IOException; import java.io.InputStreamReader; import java.util.ArrayList; public class SqlParser { public static ArrayList<String> Parser(String sqlfile, String user) { ArrayList<String> statements = new ArrayList<String>(); BufferedReader reader = null; try { reader = new BufferedReader(new InputStreamReader(new FileInputStream(sqlfile), "UTF-8")); String line = reader.readLine(); log.info("PARSE_PRINTING" + line); while (line != null) { line = line.trim(); if (line.startsWith("--") || line.equals("")) { // 跳过单行注释和空行 } else if (line.startsWith("/*")) { String s = ""; while ((s = reader.readLine()) != null && !s.endsWith("*/")) { // 跳过多行注释 } } else { if (line.endsWith(";")) { statements.add(line); } else { StringBuilder sb = new StringBuilder(line); String s = ""; while ((s = reader.readLine()) != null) { sb.append(" ").append(s); if (s.endsWith(";")) break; } statements.add(sb.toString()); System.out.println("LINE" + line); } } line = reader.readLine(); log.info("PARSE_PRINTING" + line); } int count = 0; for (String test : statements) { log.info("STATEMENT " + test); count++; } log.info("Total No Of SQL STATEMENTS " + count); return statements; } catch (Exception e) { e.printStackTrace(); } finally { if (reader != null) { try { reader.close(); } catch (IOException e) { e.printStackTrace(); } } } return statements; } }
问题排查与解决步骤
1. 确认本地文件的完整性
先在Docker容器内验证下载后的文件是否包含正确的特殊字符:
- 使用
cat命令直接查看文件内容:cat /u01/infor/sce/scp11/logs/temp.sql - 或用
hexdump查看字节值,确认ä(UTF-8字节为C3 A4)、ü(C3 BC)是否存在:hexdump -C /u01/infor/sce/scp11/logs/temp.sql | head -20
如果文件内容正确,说明问题出在Java读取后的输出/日志环节。
2. 修正Docker容器的字符编码
Java虚拟机默认继承系统编码,Docker容器默认编码可能为ASCII,导致UTF-8字符无法正常显示:
- 启动容器时添加环境变量:
docker run -e LANG=C.UTF-8 -e LC_ALL=C.UTF-8 [你的容器镜像] - 或在Java启动参数中强制指定编码:
java -Dfile.encoding=UTF-8 -Dsun.jnu.encoding=UTF-8 -jar [你的应用jar包]
3. 配置日志框架的UTF-8编码
如果使用Logback,在logback.xml中为编码器指定UTF-8:
<appender name="CONSOLE" class="ch.qos.logback.core.ConsoleAppender"> <encoder> <pattern>%d{yyyy-MM-dd HH:mm:ss} [%thread] %-5level %logger{36} - %msg%n</pattern> <charset>UTF-8</charset> </encoder> </appender>
如果使用Log4j2,在log4j2.xml中配置:
<Console name="Console" target="SYSTEM_OUT"> <PatternLayout pattern="%d{yyyy-MM-dd HH:mm:ss} [%thread] %-5level %logger{36} - %msg%n" charset="UTF-8"/> </Console>
4. 验证Java读取的字符正确性
在读取代码中添加字节验证,确认读取到的字符编码正确:
import java.nio.charset.StandardCharsets; import java.util.Arrays; // 在读取逻辑中添加 reader = new BufferedReader(new InputStreamReader(new FileInputStream(sqlfile), StandardCharsets.UTF_8)); String line = reader.readLine(); if (line != null) { // 打印字符的UTF-8字节值 byte[] utf8Bytes = line.getBytes(StandardCharsets.UTF_8); log.info("Line UTF-8 bytes: " + Arrays.toString(utf8Bytes)); }
如果字节值符合UTF-8标准(比如ä对应[195, 164]),说明读取没问题,问题出在日志输出或后续处理。
5. 避免使用默认编码的IO类
如果有写入文件的操作,禁止使用FileWriter/FileReader这类依赖系统默认编码的类,改用明确指定UTF-8的方式:
import java.nio.file.Files; import java.nio.file.Paths; import java.nio.charset.StandardCharsets; import java.io.BufferedWriter; // 写入文件示例 try (BufferedWriter writer = Files.newBufferedWriter(Paths.get(outputFilePath), StandardCharsets.UTF_8)) { writer.write(sqlContent); } catch (IOException e) { log.error("Write file failed", e); }
内容的提问来源于stack exchange,提问作者john smith
相关产品推荐
相关产品推荐

