Pandas创建含情感分析的DataFrame后出现空数据框,改列名无效
问题描述
执行以下代码后得到空DataFrame:
df=pd.DataFrame(data, columns=["Date", "Time", "Contact", "Message"]) df['Date']=pd.to_datetime(df['Date']) data=df.dropna() from nltk.sentiment.vader import SentimentIntensityAnalyzer sentiments=SentimentIntensityAnalyzer() data["positive"]=[sentiments.polarity_scores(i)["pos"] for i in data["Message"]] data["negative"]=[sentiments.polarity_scores(i)["neg"] for i in data["Message"]] data["neutral"]=[sentiments.polarity_scores(i)["neu"] for i in data["Message"]] data.head(50)
运行输出:
Empty DataFrame Columns: [Date, Time, Author, Message, Positive, Negative, Neutral] Index: []
修改列名后问题依然存在,无法生成正常数据表。
解决方案
- 检查原始
data变量有效性:创建df前先打印原始data,确认其本身是否包含有效数据,若原始数据为空,后续操作必然得到空表。 - 排查
dropna()的过度删除:df.dropna()默认删除所有含空值的行,若原始数据每行都存在空值,执行后会清空数据。可先通过df.isnull().sum()统计各列空值数,再调整参数,比如用dropna(subset=["Message"])仅删除Message列为空的行,保留其他列有缺失但Message有效的数据。 - 核对列名匹配性:输出列中出现
Author,但代码创建df时指定的是Contact,说明原始data的列名实际为Author,导致Contact列全为空,触发dropna()删除所有行。可通过print(df.columns)查看df实际列名,确保和原始数据列名一致。 - 修复日期转换异常:若
Date列格式无法被pd.to_datetime()解析,会生成NaT空值,进而被dropna()删除。可添加errors='coerce'参数:df['Date']=pd.to_datetime(df['Date'], errors='coerce'),再通过df['Date'].isnull().sum()查看转换失败的行数,针对性调整日期格式或处理异常值。
内容的提问来源于stack exchange,提问作者Siddhant Jain
相关产品推荐
相关产品推荐

