如何利用多组字典填充已创建的Pandas DataFrame?
没问题,我来帮你搞定这个DataFrame填充的需求!结合你给出的两个字典,我们可以通过遍历字典的键值对来给DataFrame赋值,下面是具体的实现步骤和代码:
1. 先初始化DataFrame(避免NaN问题)
刚创建的DataFrame里全是NaN,先把所有值填充为0,这样后续未涉及的主题(比如evolution、mathematics等)会保持0值,更符合计数逻辑:
import pandas as pd # 记得先导入pandas,如果你还没导的话 # 你的原有代码 range_of_years = range(2012, 2017) topics = ['ecology','evolution','mathematics','biogeography','neutral theory'] topic_count_timeline = pd.DataFrame(index=topics, columns=range_of_years) count_dict = {2012: 10, 2013: 20, 2014: 12, 2015: 8, 2016: 9} paper_topics_dict = {'ecology': 0.7, 'neutral theory': 0.3} # 初始化所有单元格为0 topic_count_timeline = topic_count_timeline.fillna(0)
2. 遍历字典填充数值
这里提供两种方式,你可以根据自己的需求选择:
方式一:嵌套循环(直观易懂)
先遍历paper_topics_dict里的主题和对应比例,再遍历count_dict里的年份和总计数,将总计数乘以比例后赋值给对应单元格:
# 遍历主题字典 for topic, proportion in paper_topics_dict.items(): # 遍历年份计数字典 for year, total_count in count_dict.items(): # 计算当前主题在该年份的计数并赋值 topic_count_timeline.loc[topic, year] = total_count * proportion
方式二:Pandas向量化操作(高效简洁)
把count_dict转换成Pandas Series,直接和比例相乘后赋值给整行,避免循环,处理大数据时效率更高:
# 将年份计数字典转为Series count_series = pd.Series(count_dict) # 遍历主题字典,直接赋值整行 for topic, proportion in paper_topics_dict.items(): topic_count_timeline.loc[topic] = count_series * proportion
3. 可选:转为整数计数(如果需要)
如果你的计数需要是整数,可以在赋值时加上round()或者astype(int):
# 方式一调整: topic_count_timeline.loc[topic, year] = round(total_count * proportion) # 方式二调整: topic_count_timeline.loc[topic] = (count_series * proportion).round().astype(int)
最终效果
运行后你的DataFrame会变成这样(以浮点数版本为例):
2012 2013 2014 2015 2016 ecology 7.0 14.0 8.4 5.6 6.3 evolution 0.0 0.0 0.0 0.0 0.0 mathematics 0.0 0.0 0.0 0.0 0.0 biogeography 0.0 0.0 0.0 0.0 0.0 neutral theory 3.0 6.0 3.6 2.4 2.7
内容的提问来源于stack exchange,提问作者Lorcán
相关产品推荐
相关产品推荐

