You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python 2.7读取数据库含斜撇字符串写入文件变为问号如何解决

Python2.7写入文件特殊字符变问号解决方案

问题根源

Python2中存在str(字节串)和unicode(字符串)两种类型,你遇到的问号是编码转换失败时的默认替换字符:

  1. 你从数据库取出的message为str类型的字节串,不是unicode对象
  2. 使用codecs.open写入时,模块会默认用ASCII编码将传入的str转为unicode,右单引号’不属于ASCII字符集,转换失败就被替换为问号

修复步骤

步骤1:优先从数据源侧规范编码(推荐)

连接数据库时指定编码,直接获取unicode类型的查询结果,避免后续手动转码出错。以MySQL为例,连接时新增如下参数:

import MySQLdb
# 增加charset和use_unicode参数,查询结果自动转为unicode
conn = MySQLdb.connect(
    host='你的数据库地址',
    user='用户名',
    passwd='密码',
    db='库名',
    charset='utf8',
    use_unicode=True
)

步骤2:写入前统一处理编码和字符替换

如果无法修改数据库连接配置,手动将查询结果转成unicode后再写入,同时可直接替换你需要的普通单引号:

# -*- coding: utf-8 -*-
import codecs

# 1. 先把str类型的message转成unicode,注意utf-8要和数据库实际编码一致,是gbk就改gbk
if isinstance(message, str):
    message = message.decode('utf-8')

# 2. 可选:直接把右单引号替换为你需要的普通单引号
message = message.replace(u'\u2019', u"'")

# 3. 写入文件,指定errors='strict'可在编码出错时直接抛异常,避免静默生成问号
with codecs.open(path, 'w', encoding='utf-8', errors='strict') as f:
    f.write(message)

替代写法(不用codecs模块)

普通open写入时手动把unicode编码为utf-8字节串即可:

if isinstance(message, str):
    message = message.decode('utf-8')
message = message.replace(u'\u2019', u"'")
with open(path, 'w') as f:
    f.write(message.encode('utf-8'))

内容的提问来源于stack exchange,提问作者George Hernando

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.26 02:15:03