使用pyreadstat读取SAS大文件报错:name 'row' is not defined
解决
name 'row' is not defined报错问题 错误原因
代码中调用getChunk时使用了row['/data/research1/test/cc/53/'],但row变量从未被定义,Python无法识别该变量导致报错。
修复步骤
- 替换未定义的
row变量:直接使用包含文件名的完整文件路径,把row['/data/research1/test/cc/53/']替换为拼接后的完整路径(目录路径+之前定义的filename)。 - 统一读取函数调用:循环里不要硬写
pyreadstat.read_sas7bdat,改用之前定义的getChunk函数,让代码同时兼容sas7bdat和xpt格式,避免重复逻辑。
修正后的完整代码
filename = 'df_53.SAS7BDAT' CHUNKSIZE = 50000 offset = 0 # 统一获取读取函数 if filename.lower().endswith('sas7bdat'): getChunk = pyreadstat.read_sas7bdat else: getChunk = pyreadstat.read_xport # 拼接完整文件路径,替换未定义的row变量 full_path = f'/data/research1/test/cc/53/{filename}' allChunk,_ = getChunk(full_path, row_limit=CHUNKSIZE, row_offset=offset) allChunk = allChunk.astype('category') while True: offset += CHUNKSIZE # 统一用getChunk读取,兼容两种格式 chunk, _ = getChunk(full_path, row_limit=CHUNKSIZE, row_offset=offset) if chunk.empty: break # 空块说明数据已读完,终止循环 for eachCol in chunk: # 合并分类变量的类别,保证一致性 colUnion = pd.api.types.union_categoricals([allChunk[eachCol], chunk[eachCol]]) allChunk[eachCol] = pd.Categorical(allChunk[eachCol], categories=colUnion.categories) chunk[eachCol] = pd.Categorical(chunk[eachCol], categories=colUnion.categories) allChunk = pd.concat([allChunk, chunk]) # 追加块到结果DataFrame
内容的提问来源于stack exchange,提问作者lydias
相关产品推荐
相关产品推荐

