使用Google Cloud Datastore Python客户端时整数转浮点数问题排查
我用google-cloud-datastore Python模块编写了脚本,将CSV数据上传至Cloud Datastore。脚本运行正常,但发现CSV中的整数值被存储为浮点数。请问这是数据上传至Datastore的默认行为,还是我的代码存在问题?
以下是我的代码:
import sys import getopt import pandas as pd from google.cloud import datastore def write_dict_chunks(data, SIZE=100): log_count = 0 datastore_client = datastore.Client() task_key = datastore_client.key(kind) for i in xrange(0, len(data), SIZE): entities = [] for each_entry in data[i : i+SIZE]: nan_check = lambda v: v if str(v)!='nan' else None string_check = lambda v: v.decode('utf-8') if isinstance(v, str) else v write_row = {k: nan_check(string_check(v)) for k, v in each_entry.iteritems()} entity = datastore.Entity(key=task_key) entity.update(write_row) entities.append(entity) datastore_client.put_multi(entities) log_count += len(entities) print 'Wrote {} entities to datastore'.format(log_count) try: opts, args = getopt.getopt(sys.argv[1:], "ho:v", ["kind=", "filepath="]) if len(args) > 0: for each in args: print 'Unrecognized argument: '+each sys.exit(2) except getopt.GetoptError as err: # print help information and exit: print str(err) # will print something like "option -a not recognized" print 'Usage: python parse_csv.py --kind=kind_name --filepath=path_to_csv' kind = None filepath = None for option, argument in opts: if option in '--kind': kind = argument elif option in '--filepath': filepath = argument df = pd.read_csv(filepath) df = df.to_dict(orient='records') write_dict_chunks(df)
这不是Cloud Datastore的默认行为——Datastore本身完全支持整数(int)类型存储,问题出在Pandas读取CSV时的数据类型转换以及你的代码处理环节。
问题根源
当你用pd.read_csv(filepath)读取CSV时,如果整数列中存在缺失值(比如空单元格),Pandas会自动把该列的类型从int转为float——因为Python原生int类型无法表示缺失值(NaN是float类型的特殊值)。后续你把DataFrame转成字典列表后,这些原本的整数值已经是float类型了,上传到Datastore自然会被存储为浮点数。
另外,你的代码里的nan_check只是把NaN转为None,但并没有把可以还原为整数的float值(比如10.0)转回int类型,这就导致最终上传的数据保留了浮点类型。
解决方案
你可以通过以下几种方式修复这个问题:
1. 读取CSV时指定整数列的类型(适用于无缺失值的列)
如果你的整数列没有缺失值,可以在读取CSV时直接指定dtype参数,强制该列为整数类型:
# 假设CSV中有"id"和"age"两个整数列 df = pd.read_csv(filepath, dtype={"id": int, "age": int})
2. 处理字典时自动转换可还原的浮点数为整数
如果列中存在缺失值,你可以在生成write_row时,增加一步类型转换:把值为整数形式的浮点数(比如5.0)转成int,同时保留真正的浮点数(比如5.5):
def convert_float_to_int(v): # 如果是浮点数且是整数形式,转为int;否则保持原类型 if isinstance(v, float) and v.is_integer(): return int(v) return v # 在生成write_row时加入这个转换 write_row = { k: nan_check(string_check(convert_float_to_int(v))) for k, v in each_entry.iteritems() }
3. 用Int64类型处理带缺失值的整数列(推荐)
Pandas支持Int64(注意大写I)类型,它可以同时存储整数和缺失值(用pd.NA表示)。你可以在读取CSV时指定该类型:
df = pd.read_csv(filepath, dtype={"id": "Int64", "age": "Int64"})
转换为字典后,缺失值会是pd.NA,你只需在nan_check里把pd.NA也转为None,整数就能保持int类型。
验证
修改后,你可以在生成write_row时打印值的类型,确认整数已经被正确处理,再上传到Datastore,就能看到数据以整数类型存储了。
内容的提问来源于stack exchange,提问作者Tameem

