Python pandas中Object转Int64失败求助(astype/to_numeric无效)
解决方法
问题出在你的数据里有带千分位逗号的字符串(比如'4,600'),直接转int64会因为逗号不是有效数字字符报错。得先把逗号去掉再转换,或者在读取CSV的时候就处理好。
方案1:读取CSV时直接处理(推荐)
修改loadDataFile方法,利用pandas.read_csv的thousands参数自动解析带逗号的数值,同时用na_values把'na'标记为缺失值,之后再填充为0:
def loadDataFile(self): # thousands=',' 自动识别千分位逗号,na_values='na'把na设为NaN self.df = pd.read_csv('Int_Monthly_Visitor.csv', index_col=0, thousands=',', na_values='na') # 把NaN填充为0,转成int64类型 self.df = self.df.fillna(0).astype(np.int64)
这样后续parseData里直接取place就已经是int64类型了,不需要再转换:
def parseData(self): location = self.area[self.areaRegion] place = self.df.filter(items=location) print(place.info())
方案2:在parseData里处理现有数据
如果不想修改读取逻辑,就在parseData里先清理逗号再转类型:
def parseData(self): location = self.area[self.areaRegion] place = self.df.filter(items=location) # 先替换逗号,填充缺失值后转int64 place = place.apply(lambda col: col.str.replace(',', '').fillna('0').astype(np.int64)) print(place.info())
为什么之前的方法失败?
不管是astype、to_numeric还是指定dtype,都无法直接识别带逗号的字符串为数字,必须先移除逗号这类非数字符号,才能完成类型转换。
内容的提问来源于stack exchange,提问作者Noob that needs help
相关产品推荐
相关产品推荐

