You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

.dta转.xlsx遇字符编码错误,求解决方案(附Pandas代码)

解决Pandas读取越南语编码Stata .dta文件的编码错误问题

方法1:指定越南语编码直接读取

越南语常用编码为cp1258,直接在read_stata中指定该编码,可规避默认utf-8解码失败的问题:

import pandas as pd

# 使用cp1258编码读取文件
df = pd.read_stata("D:\STATATATATA\ho1_20.dta", encoding="cp1258")
df.to_excel('file.xlsx', index=False)

如果cp1258无效,可先以latin-1读取再转码修正:

import pandas as pd

df = pd.read_stata("D:\STATATATATA\ho1_20.dta", encoding="latin-1")
# 遍历所有字符串列,转码为正确的越南语编码
for col in df.select_dtypes(include=['object']).columns:
    df[col] = df[col].str.encode('latin-1').str.decode('cp1258')
df.to_excel('file.xlsx', index=False)

方法2:用Stata直接转码导出

如果Python端处理麻烦,直接在Stata中完成转码和导出更稳妥:

  1. 打开Stata加载目标.dta文件:
    use "D:\STATATATATA\ho1_20.dta", clear
    
  2. 设置原文件编码为越南语cp1258并转成utf-8格式:
    unicode encoding set cp1258
    unicode translate "D:\STATATATATA\ho1_20_utf8.dta", transutf8
    
  3. 导出为Excel文件:
    export excel using "file.xlsx", firstrow(variables) replace
    

方法3:检测文件编码后精准读取

若不确定文件具体编码,用chardet库检测后再读取:

import chardet
import pandas as pd

# 检测文件编码
with open("D:\STATATATATA\ho1_20.dta", 'rb') as f:
    detect_result = chardet.detect(f.read())
target_encoding = detect_result['encoding']

# 用检测到的编码读取文件
df = pd.read_stata("D:\STATATATATA\ho1_20.dta", encoding=target_encoding)
df.to_excel('file.xlsx', index=False)

若检测结果为None,可尝试cp1258、utf-16这类越南语常用编码逐一测试。

内容的提问来源于stack exchange,提问作者kennyS

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 15:07:55