为何pandas.to_datetime()无法更改DataFrame列的数据类型?
原问题
尝试将DataFrame中的time列转换为带时区信息的datetime64类型,但转换未生效:
print(f'1:\n{filtered_df.tail(1)}\ndtype: {filtered_df["time"].dtypes}') filtered_df["time"] = pandas.to_datetime(filtered_df["time"]) print(f'2:\n{filtered_df.tail(1)}\ndtype: {filtered_df["time"].dtypes}')
输出结果:
1: id symbol time open high low close volume 23 75979 USDCAD 2022-11-06 19:11:00-05:00 1.35102 1.35113 1.35102 1.35114 None dtype: object 2: id symbol time open high low close volume 23 75979 USDCAD 2022-11-06 19:11:00-05:00 1.35102 1.35113 1.35102 1.35114 None dtype: object
但在新创建的DataFrame中执行相同操作时,转换成功:
df = pandas.DataFrame({"id": [75979], "symbol": ["USDCAD"], "time": ["2022-11-06 19:11:00-05:00"], "open": [1.35102], "high": [1.3513], "low": [1.35102], "close": [1.35114], "volume": [None]}) print(f'pre: {df["time"].dtypes}') df["time"] = pandas.to_datetime(df["time"]) print(f'post: {df["time"].dtypes}') print(df)
输出结果:
pre: object post: datetime64[ns, pytz.FixedOffset(-300)] id symbol time open high low close volume 0 75979 USDCAD 2022-11-06 19:11:00-05:00 1.35102 1.3513 1.35102 1.35114 None
无法理解两种场景结果不同的原因。
更新
原本以为原问题中的最小可复现示例(MRE)足够,但实际并非如此。对代码大量修改后问题已解决,但回头创建可用的MRE却失败了。
简言之,不清楚究竟是什么操作导致了最初的问题。最初是通过.concat()函数生成filtered_df的基础数据,但这似乎不是问题根源。
目前问题已解决,暂时选择保留问题而非删除。以下是尝试复现原问题的失败代码,仅作背景参考:
import pandas df1 = pandas.DataFrame({"id": [75979], "symbol": ["USDCAD"], "time": ["2022-11-06 19:11:00-05:00"], "open": [1], "high": [1], "low": [1], "close": [1], "volume": [None]}) df2 = pandas.DataFrame({"id": [75980], "symbol": ["USDCAD"], "time": ["2022-11-06 19:12:00-05:00"], "open": [2], "high": [2], "low": [2], "close": [2], "volume": [None]}) df = pandas.concat([df1, df2]) df.reset_index(drop=True, inplace=True) df
输出结果:
id symbol time open high low close volume 0 75979 USDCAD 2022-11-06 19:11:00-05:00 1 1 1 1 None 1 75980 USDCAD 2022-11-06 19:12:00-05:00 2 2 2 2 None
df.info()
输出结果:
<class 'pandas.core.frame.DataFrame'> RangeIndex: 2 entries, 0 to 1 Data columns (total 8 columns): # Column Non-Null Count Dtype --- ------ -------------- ----- 0 id 2 non-null int64 1 symbol 2 non-null object 2 time 2 non-null object 3 open 2 non-null int64 4 high 2 non-null int64 5 low 2 non-null int64 6 close 2 non-null int64 7 volume 0 non-null object dtypes: int64(5), object(3) memory usage: 256.0+ bytes
df["time"] = pandas.to_datetime(df["time"]) df
输出结果:
id symbol time open high low close volume 0 75979 USDCAD 2022-11-06 19:11:00-05:00 1 1 1 1 None 1 75980 USDCAD 2022-11-06 19:12:00-05:00 2 2 2 2 None id symbol time open high low close volume 0 75979 USDCAD 2022-11-06 19:11:00-05:00 1 1 1 1 None 1 75980 USDCAD 2022-11-06 19:12:00-05:00 2 2 2 2 None
df.info()
输出结果:
<class 'pandas.core.frame.DataFrame'> RangeIndex: 2 entries, 0 to 1 Data columns (total 8 columns): # Column Non-Null Count Dtype --- ------ -------------- ----- 0 id 2 non-null int64 1 symbol 2 non-null object 2 time 2 non-null datetime64[ns, pytz.FixedOffset(-300)] 3 open 2 non-null int64 4 high 2 non-null int64 5 low 2 non-null int64 6 close 2 non-null int64 7 volume 0 non-null object dtypes: datetime64[ns, pytz.FixedOffset(-300)](1), int64(5), object(2) memory usage: 256.0+ bytes
内容的提问来源于stack exchange,提问作者Jason
相关产品推荐
相关产品推荐

