You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Java读取S3下载文件时特殊字符丢失问题求助

S3文件特殊字符(ä、ü)读取后变为空格的解决方法

S3存储的文件包含ä、ü这类特殊字符,Java下载后读取/写入时这些字符会被替换为空空格。尝试过多种UTF-8格式读写方式(BufferedReader、BufferedWriter、FileReader、FileWriter等)均无效,且已确认S3下载环节无问题,问题出在Java读取后的处理流程中。

当前代码片段

Java下载代码

String uploadPath = "/u01/infor/sce/scp11/logs/temp.sql";
AmazonS3 s3Client = AmazonS3ClientBuilder.defaultClient();
S3Object obj = s3Client.getObject(System.getenv("BUCKET"), "/temp.sql");
S3ObjectInputStream s3ObjectInputStream = obj.response().asInputStream();

// 下载到本地文件
String localFilePath = uploadPath;
try (OutputStream outputStream = new FileOutputStream(localFilePath)) {
    byte[] buffer = new byte[1024];
    int bytesRead;
    while ((bytesRead = s3ObjectInputStream.read(buffer)) != -1) {
        outputStream.write(buffer, 0, bytesRead);
    }

    System.out.println("File downloaded successfully.");
} catch (Exception e) {
    log.info("Error while downloading sql file from s3" + e.getMessage(), e);
}

可正常处理特殊字符的Python代码

import io
import json
import boto3
import botocore
import datetime
import base64
import os
import random
import string
import time


s3 = boto3.resource('s3')

s3_object = boto3.resource('s3').Object('hg-us-east-1', 'logs/delete-entries.sql')
data = io.BytesIO()
s3_object.download_fileobj(data)
s = None
b = s3_object.get()['Body']
s = b.read().decode('utf-8')
print(s)
# 写入时明确指定UTF-8编码,避免依赖系统默认
f = open("c:\\Users\\tim\\test.sql", "w", encoding='utf-8')
f.write(s)
f.close()

Java读取解析代码

import java.io.BufferedReader;
import java.io.FileInputStream;
import java.io.IOException;
import java.io.InputStreamReader;
import java.util.ArrayList;

public class SqlParser {
    public static ArrayList<String> Parser(String sqlfile, String user) {
        ArrayList<String> statements = new ArrayList<String>();
        BufferedReader reader = null;
        try {
            reader = new BufferedReader(new InputStreamReader(new FileInputStream(sqlfile), "UTF-8"));
            String line = reader.readLine();
            log.info("PARSE_PRINTING" + line);
            while (line != null) {
                line = line.trim();
                if (line.startsWith("--") || line.equals("")) {
                    // 跳过单行注释和空行
                } else if (line.startsWith("/*")) {
                    String s = "";
                    while ((s = reader.readLine()) != null && !s.endsWith("*/")) {
                        // 跳过多行注释
                    }
                } else {
                    if (line.endsWith(";")) {
                        statements.add(line);
                    } else {
                        StringBuilder sb = new StringBuilder(line);
                        String s = "";
                        while ((s = reader.readLine()) != null) {
                            sb.append(" ").append(s);
                            if (s.endsWith(";"))
                                break;
                        }
                        statements.add(sb.toString());
                        System.out.println("LINE" + line);
                    }
                }
                line = reader.readLine();
                log.info("PARSE_PRINTING" + line);
            }
            int count = 0;
            for (String test : statements) {
                log.info("STATEMENT " + test);
                count++;
            }
            log.info("Total No Of SQL STATEMENTS " + count);
            return statements;
        } catch (Exception e) {
            e.printStackTrace();
        } finally {
            if (reader != null) {
                try {
                    reader.close();
                } catch (IOException e) {
                    e.printStackTrace();
                }
            }
        }
        return statements;
    }
}

问题排查与解决步骤

1. 确认本地文件的完整性

先在Docker容器内验证下载后的文件是否包含正确的特殊字符:

  • 使用cat命令直接查看文件内容:
    cat /u01/infor/sce/scp11/logs/temp.sql
    
  • 或用hexdump查看字节值,确认ä(UTF-8字节为C3 A4)、ü(C3 BC)是否存在:
    hexdump -C /u01/infor/sce/scp11/logs/temp.sql | head -20
    

如果文件内容正确,说明问题出在Java读取后的输出/日志环节。

2. 修正Docker容器的字符编码

Java虚拟机默认继承系统编码,Docker容器默认编码可能为ASCII,导致UTF-8字符无法正常显示:

  • 启动容器时添加环境变量:
    docker run -e LANG=C.UTF-8 -e LC_ALL=C.UTF-8 [你的容器镜像]
    
  • 或在Java启动参数中强制指定编码:
    java -Dfile.encoding=UTF-8 -Dsun.jnu.encoding=UTF-8 -jar [你的应用jar包]
    

3. 配置日志框架的UTF-8编码

如果使用Logback,在logback.xml中为编码器指定UTF-8:

<appender name="CONSOLE" class="ch.qos.logback.core.ConsoleAppender">
  <encoder>
    <pattern>%d{yyyy-MM-dd HH:mm:ss} [%thread] %-5level %logger{36} - %msg%n</pattern>
    <charset>UTF-8</charset>
  </encoder>
</appender>

如果使用Log4j2,在log4j2.xml中配置:

<Console name="Console" target="SYSTEM_OUT">
  <PatternLayout pattern="%d{yyyy-MM-dd HH:mm:ss} [%thread] %-5level %logger{36} - %msg%n" charset="UTF-8"/>
</Console>

4. 验证Java读取的字符正确性

在读取代码中添加字节验证,确认读取到的字符编码正确:

import java.nio.charset.StandardCharsets;
import java.util.Arrays;

// 在读取逻辑中添加
reader = new BufferedReader(new InputStreamReader(new FileInputStream(sqlfile), StandardCharsets.UTF_8));
String line = reader.readLine();
if (line != null) {
    // 打印字符的UTF-8字节值
    byte[] utf8Bytes = line.getBytes(StandardCharsets.UTF_8);
    log.info("Line UTF-8 bytes: " + Arrays.toString(utf8Bytes));
}

如果字节值符合UTF-8标准(比如ä对应[195, 164]),说明读取没问题,问题出在日志输出或后续处理。

5. 避免使用默认编码的IO类

如果有写入文件的操作,禁止使用FileWriter/FileReader这类依赖系统默认编码的类,改用明确指定UTF-8的方式:

import java.nio.file.Files;
import java.nio.file.Paths;
import java.nio.charset.StandardCharsets;
import java.io.BufferedWriter;

// 写入文件示例
try (BufferedWriter writer = Files.newBufferedWriter(Paths.get(outputFilePath), StandardCharsets.UTF_8)) {
    writer.write(sqlContent);
} catch (IOException e) {
    log.error("Write file failed", e);
}

内容的提问来源于stack exchange,提问作者john smith

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.15 04:47:08