Windows与Linux环境下MySQL5数据库波兰字符编码乱码问题求助
你已经做了不少基础配置工作,但还是踩了字符编码的常见坑,我来分享几个针对性的排查方向和解决方案:
1. 检查MySQL服务器的全局字符集配置
Linux上的MySQL经常会出现「数据库/表排序规则正确,但服务器全局字符集不匹配」的情况,这会直接影响数据的存储编码。
登录远程MySQL终端,执行以下命令查看全局字符集参数:
SHOW VARIABLES LIKE 'character_set%'; SHOW VARIABLES LIKE 'collation%';重点确认以下参数的值为
utf8(更推荐用utf8mb4,支持全Unicode字符):character_set_servercharacter_set_clientcharacter_set_connectioncharacter_set_results
如果参数不符合预期,修改MySQL配置文件(Ubuntu上通常是
/etc/mysql/mysql.conf.d/mysqld.cnf),添加或更新配置:[mysqld] character-set-server=utf8mb4 collation-server=utf8mb4_unicode_ci [client] default-character-set=utf8mb4保存后重启MySQL服务:
sudo systemctl restart mysql
2. 优化JDBC连接字符串的编码参数
你当前的连接字符串用了characterEncoding=utf-8,MySQL JDBC驱动对小写的utf-8处理可能存在兼容问题,建议换成规范的UTF-8,同时补充几个关键参数:
jdbc:mysql://localhost/dbname?autoReconnect=true&useUnicode=true&characterEncoding=UTF-8&createDatabaseIfNotExist=true&useSSL=false&useLegacyDatetimeCode=false
3. 调整Hibernate的方言与编码配置
MySQL5InnoDBDialect对utf8mb4的支持不够完善,如果你的MySQL版本是5.7及以上,建议替换为对应方言,同时统一编码配置:
@Bean public SessionFactory sessionFactory() throws Exception { LocalSessionFactoryBean sessionFactoryBean = new LocalSessionFactoryBean(); sessionFactoryBean.setDataSource(dataSource()); sessionFactoryBean.setPackagesToScan(new String[] { "x.domain" }); Properties hibernateProperties = new Properties(); hibernateProperties.put("hibernate.show_sql", false); hibernateProperties.put("hibernate.bytecode.use_reflection_optimizer", false); hibernateProperties.put("hibernate.check_nullability", false); // 替换为适配utf8mb4的方言 hibernateProperties.put("hibernate.dialect", "org.hibernate.dialect.MySQL57InnoDBDialect"); hibernateProperties.put("hibernate.search.autoregister_listeners", false); // 统一使用utf8mb4编码 hibernateProperties.put("hibernate.connection.CharSet", "utf8mb4"); hibernateProperties.put("hibernate.connection.characterEncoding", "utf8mb4"); hibernateProperties.put("hibernate.connection.useUnicode", true); sessionFactoryBean.setHibernateProperties(hibernateProperties); sessionFactoryBean.afterPropertiesSet(); return sessionFactoryBean.getObject(); }
4. 验证数据解析阶段的编码
有时候问题出在jsoup解析环节,你可以在Linux服务器上临时打印解析后的字符串(比如System.out.println(parsedText)),如果控制台显示也是问号,说明解析时编码不匹配。需要明确指定jsoup的解析编码:
// 根据目标网页的实际编码调整,比如UTF-8或ISO-8859-2 Document doc = Jsoup.connect(url).charset("UTF-8").get();
5. 检查Linux系统的Locale设置
MySQL服务启动时会继承系统的Locale,如果系统默认不是UTF-8,也会干扰字符处理。执行以下命令查看当前Locale:
locale
确保LANG和LC_ALL的值是en_US.UTF-8或pl_PL.UTF-8这类UTF-8格式的Locale。如果不是,执行以下命令配置(Ubuntu为例):
sudo locale-gen pl_PL.UTF-8 sudo update-locale LANG=pl_PL.UTF-8 LC_ALL=pl_PL.UTF-8
配置完成后重启MySQL服务,必要时重启服务器确保生效。
建议先从全局字符集配置开始排查,这是Linux环境下字符编码问题最常见的根源。
内容的提问来源于stack exchange,提问作者Sarbitar

