You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何统计DataFrame指定值出现频次并按年代维度绘制折线图

解决方案

你之前的代码问题在于直接按单个出版年份分组计数,没有做十年段的维度聚合,所以输出结果和按十年统计的预期不符,按以下步骤实现即可:

核心逻辑

  • 先筛选出目标城市的所有相关记录,排除其他城市的干扰数据
  • 将每条记录的出版年份向下取整到对应十年段的起始年(例如1854、1859都归为1850段,1866归为1860段)
  • 按十年段分组统计提及数量,可选补全计数为0的十年段避免折线断点
  • 配置绘图参数输出符合预期的折线图

示例数据结构参考

你提供的测试DataFrame结构如下:

title of the novel                author          publishing year   mentioned cities   
0   Beasts and creatures        Bruno Ivory             1850           New York 
1   Monsters                    Renata Mcniar           1866           New York 
2   At risk                     Charles Dobi            1870           New York   
3   Manuela and Ricardo         Lucas Zacci             1889           New York
4   War against the machine     Angelina Trotter        1854           New York

完整实现代码

import pandas as pd
import matplotlib.pyplot as plt

# 替换为你实际的数据集读取逻辑
# df = pd.read_csv("your_novel_data_source.csv")

# 配置统计目标
target_city = "New York"

# 第一步:筛选目标城市的所有记录
target_city_df = df[df["mentioned cities"] == target_city].copy()

# 第二步:生成出版十年段字段
target_city_df["decade"] = (target_city_df["publishing year"] // 10) * 10

# 第三步:按十年段聚合统计提及次数
decade_count_result = target_city_df["decade"].value_counts().sort_index().reset_index()
decade_count_result.columns = ["publishing_decade", "novel_count"]

# 可选:补全统计区间内无记录的十年段,填充计数为0,避免折线出现断档
min_decade_val = decade_count_result["publishing_decade"].min()
max_decade_val = decade_count_result["publishing_decade"].max()
full_decade_range = range(min_decade_val, max_decade_val + 10, 10)
decade_count_result = decade_count_result.set_index("publishing_decade").reindex(full_decade_range, fill_value=0).reset_index()

# 第四步:绘制折线图
plt.figure(figsize=(10, 6))
plt.plot(
    decade_count_result["publishing_decade"],
    decade_count_result["novel_count"],
    marker="o",
    linestyle="-",
    linewidth=2
)
plt.title(f"Mention frequency of {target_city} in novels by publishing decade")
plt.xlabel("Publishing decade")
plt.ylabel("Number of novels mentioning the city")
plt.grid(linestyle="--", alpha=0.3)
plt.xticks(full_decade_range, rotation=45)
plt.tight_layout()
plt.show()

预期效果参考

你在Excel中制作的效果参考如下:
十年段统计折线图预期效果

原有代码的偏差说明

你之前尝试的代码存在两个核心问题,导致结果不符合预期:

  • 分组统计的维度是单个出版年份,没有做十年段的转换聚合,统计粒度不匹配需求
  • 没有处理无任何提及记录的十年段,会导致横轴刻度不连续,折线出现不必要的缺口

内容的提问来源于stack exchange,提问作者Digital_humanities

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.27 05:15:31