使用PyArrow读取CSV时,\N空值转换int64报错的解决求助
解决PyArrow读取CSV时"\N"无法识别为空值的问题
问题出在ConvertOptions的null_values参数设置上:你当前设为null_values=True,这并不是指定自定义空值标记的正确方式。PyArrow默认不会将"\N"识别为空值,所以需要手动把这个字符串加入空值列表。
修改方案
将null_values设置为包含"\N"的字符串列表,同时可以保留PyArrow默认的空值标记(比如空字符串、"NaN"等),确保其他空值情况也能正常处理:
parse_options = csv.ParseOptions(delimiter=chr(1)) read_options = csv.ReadOptions(column_names=columns) # 自定义空值列表,加入"\N",同时保留默认空值 convert_options = csv.ConvertOptions( column_types=schema_table, include_columns=columns, include_missing_columns=True, null_values=["", "NaN", "n/a", "\N"] # 这里添加"\N" ) with hdfs.open_input_file("path") as f: csv_file = csv.read_csv(f, read_options=read_options, parse_options=parse_options, convert_options=convert_options)
说明
null_values参数接受字符串或字符串列表,用来指定哪些值应该被解析为null。- 如果只需要把
"\N"当作空值,也可以直接设为null_values="\N",但建议保留默认的空值标记,避免其他空值场景出错。 - 这样修改后,PyArrow会将字段中的
"\N"识别为null,就能正常转换为对应的int64类型了。
内容的提问来源于stack exchange,提问作者maxgm
相关产品推荐
相关产品推荐

