You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用pyreadstat读取SAS大文件报错:name 'row' is not defined

解决name 'row' is not defined报错问题

错误原因

代码中调用getChunk时使用了row['/data/research1/test/cc/53/'],但row变量从未被定义,Python无法识别该变量导致报错。

修复步骤

  1. 替换未定义的row变量:直接使用包含文件名的完整文件路径,把row['/data/research1/test/cc/53/']替换为拼接后的完整路径(目录路径+之前定义的filename)。
  2. 统一读取函数调用:循环里不要硬写pyreadstat.read_sas7bdat,改用之前定义的getChunk函数,让代码同时兼容sas7bdat和xpt格式,避免重复逻辑。

修正后的完整代码

filename = 'df_53.SAS7BDAT'
CHUNKSIZE = 50000
offset = 0
# 统一获取读取函数
if filename.lower().endswith('sas7bdat'):
    getChunk = pyreadstat.read_sas7bdat
else:
    getChunk = pyreadstat.read_xport
# 拼接完整文件路径,替换未定义的row变量
full_path = f'/data/research1/test/cc/53/{filename}'
allChunk,_ = getChunk(full_path, row_limit=CHUNKSIZE, row_offset=offset)
allChunk = allChunk.astype('category')

while True:
    offset += CHUNKSIZE
    # 统一用getChunk读取,兼容两种格式
    chunk, _ = getChunk(full_path, row_limit=CHUNKSIZE, row_offset=offset)
    if chunk.empty: break  # 空块说明数据已读完,终止循环

    for eachCol in chunk:  # 合并分类变量的类别,保证一致性
        colUnion = pd.api.types.union_categoricals([allChunk[eachCol], chunk[eachCol]])
        allChunk[eachCol] = pd.Categorical(allChunk[eachCol], categories=colUnion.categories)
        chunk[eachCol] = pd.Categorical(chunk[eachCol], categories=colUnion.categories)

    allChunk = pd.concat([allChunk, chunk])  # 追加块到结果DataFrame

内容的提问来源于stack exchange,提问作者lydias

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 16:20:23