IBGE微数据脚本运行报错:无法下载文档及处理固定宽度文件
解决IBGE人口普查微数据下载及处理报错问题
问题场景
运行以下Python脚本调用ibgeparser库获取巴西IBGE人口普查微数据:
from ibgeparser.microdados import Microdados from ibgeparser.enums import Anos, Estados, Modalidades if __name__ == "__main__": ano = Anos.DEZ estados = [Estados.ACRE] modalidades = [Modalidades.DOMICILIOS] ibgeparser = Microdados() ibgeparser.obter_dados_ibge(ano, estados, modalidades)
报错信息
E 240105 02:22:56 log:24] Error downloading documentation data: An error occurred while trying to download the documentation data. [E 240105 02:22:56 log:24] Error extracting data from desired file: An error occurred while trying to extract data from the desired file. [E 240105 02:22:56 log:24] Error accessing documentation data: An error occurred while trying to access the documentation data. [I 240105 02:22:56 log:21] Downloading information from the state of Acre: Downloading information from the state of Acre. [E 240105 02:25:07 log:24] Error downloading documentation data: An error occurred while trying to download the documentation data. [E 240105 02:25:07 log:24] Error extracting data from desired file: An error occurred while trying to extract data from the desired file. [E 240105 02:25:07 log:24] Error processing documentation data: An error occurred while processing the documentation data.
问题根源
- URL路径错误:原脚本使用的FTP路径占位符无法匹配IBGE实际目录结构,2010年普查数据的正确路径为
https://ftp.ibge.gov.br/Censos/Censo_Demografico_2010/Resultados_Gerais_da_Amostra/Microdados/。 - 枚举与文件命名不匹配:
Modalidades.DOMICILIOS的枚举值可能与Excel工作表名称、微数据文件名不一致,导致无法找到对应资源或键。 - 错误日志信息模糊:原脚本未输出具体异常细节,难以精准定位问题。
解决方案
1. 修正数据下载URL
将原脚本中的URL常量替换为正确的HTTPS路径(IBGE支持HTTPS访问,比FTP更稳定):
# 适配2010年普查数据的固定路径 URL_DOCUMENTACAO='https://ftp.ibge.gov.br/Censos/Censo_Demografico_2010/Resultados_Gerais_da_Amostra/Microdados/Documentacao.zip' URL_MICRODADOS='https://ftp.ibge.gov.br/Censos/Censo_Demografico_2010/Resultados_Gerais_da_Amostra/Microdados/{}.zip'
若需支持多年份,调整Anos枚举的value为对应目录名(如Anos.DEZ = (10, "2010")),再修改URL为:
URL_DOCUMENTACAO='https://ftp.ibge.gov.br/Censos/Censo_Demografico_{}/Resultados_Gerais_da_Amostra/Microdados/Documentacao.zip' URL_MICRODADOS='https://ftp.ibge.gov.br/Censos/Censo_Demografico_{}/Resultados_Gerais_da_Amostra/Microdados/{}.zip'
2. 对齐枚举值与文件命名规则
检查Modalidades枚举的value,确保与IBGE文件规则完全匹配:
- 确保
Modalidades.DOMICILIOS的valor_modalidade与Layout_microdados_Amostra.xls中的工作表名称一致(如工作表名为Domicilios,则枚举值应为("Domicilios", "Domicilios"))。 - 微数据文件名格式为
Amostra_{}_{}.txt,第一个占位符需与枚举的descricao_modalidade一致(如Domicilios对应文件名Amostra_Domicilios_1.txt,ACRE的valor_estado为1)。
3. 增强错误日志排查能力
修改__download_arquivo和__extrair_arquivo方法的错误日志,输出具体异常信息:
def __download_arquivo(self, url:str, destino:str): try: with urllib.request.urlopen(url) as response, open(destino, 'wb') as out_file: shutil.copyfileobj(response, out_file) return destino except Exception as e: log.error('Error in downloading the documentation: {}'.format(str(e))) def __extrair_arquivo(self, caminho_origem:str, nome_arquivo:str, caminho_destino:str): try: with ZipFile(caminho_origem, 'r') as zf: log.debug('Extraindo o arquivo {}'.format(caminho_origem)) for membro in zf.namelist(): if os.path.basename(membro) == nome_arquivo: arquivo = zf.open(membro) caminho_arquivo = os.path.join(caminho_destino, nome_arquivo) destino = open(caminho_arquivo, "wb") with arquivo, destino: shutil.copyfileobj(arquivo, destino) return caminho_arquivo except Exception as e: log.error('Error at extracting data from desired file: {}'.format(str(e)))
4. 手动验证文件存在性
访问修正后的URL,确认Documentacao.zip和对应州的微数据zip文件(如AC.zip)是否存在,避免因IBGE文件结构变动导致路径失效。
内容的提问来源于stack exchange,提问作者andrellima
相关产品推荐
相关产品推荐

