使用Python将DataFrame写入CSV文件时遭遇编码错误
解决DataFrame写入CSV时的'charmap'编码错误
你遇到的问题核心在于提前打开文件时未指定编码,导致to_csv里的encoding参数被忽略。当你用open()打开文件却不指定编码时,系统会使用默认的charmap编码(Windows下通常是cp1252),后续to_csv传入这个已打开的文件对象时,会直接沿用文件对象的编码,你指定的encoding='cp1252'等参数根本不会生效——这就是换了多种编码仍报错的原因。
方案1:让pandas直接处理文件路径(推荐)
不用手动打开文件,直接把文件路径传给to_csv,此时encoding参数会正常生效:
try: print("Connecting to Azure and downloading data...") cnxn = pyodbc.connect('DRIVER={ODBC Driver 17 for SQL Server};SERVER='+server+';DATABASE='+database+';ENCRYPT=yes;UID='+username+';Authentication=ActiveDirectoryInteractive') with open(sqlFileName, 'r') as sql: selectStatement = sql.read() # Read result of sql into DataFrame dfAzure = pd.read_sql(selectStatement,cnxn) print("Data has been downloaded to DataFrame") # 直接传入文件路径,指定编码 dfAzure.to_csv(finalOutputFileAzure, encoding='utf-8', sep="|", index=False, errors="replace") print("Finished downloading data from Azure!") return 0 except Exception as e: print("Error message: " + str(e)) return -1
方案2:手动打开文件时指定编码
如果一定要手动打开文件,打开时就指定正确的编码,和to_csv的编码保持一致(或者to_csv里可以省略encoding参数):
try: print("Connecting to Azure and downloading data...") cnxn = pyodbc.connect('DRIVER={ODBC Driver 17 for SQL Server};SERVER='+server+';DATABASE='+database+';ENCRYPT=yes;UID='+username+';Authentication=ActiveDirectoryInteractive') # 打开文件时指定编码,比如utf-8 outfile = open(finalOutputFileAzure, "w", encoding='utf-8') with open(sqlFileName, 'r') as sql: selectStatement = sql.read() # Read result of sql into DataFrame dfAzure = pd.read_sql(selectStatement,cnxn) print("Data has been downloaded to DataFrame") # 这里可以不用再指定encoding,或者和open时保持一致 dfAzure.to_csv(outfile, sep="|", index=False, errors="replace") outfile.close() print("Finished downloading data from Azure!") return 0 except Exception as e: print("Error message: " + str(e)) return -1
另外,\ufffd是替换字符(通常表示原字符无法被编码识别),你设置的errors="replace"可以把无法编码的字符替换成?,避免报错——这个参数是有效的,但前提是编码参数能正常生效。
内容的提问来源于stack exchange,提问作者S. Hasan
相关产品推荐
相关产品推荐

