如何按年份区间分组并基于PctAgreeUS子集化unvoting.csv数据?
按年份区间分组并分析PctAgreeUS的实现方法
R语言方案
- 先加载数据:
# 读取csv文件 df <- read.csv("unvoting.csv") - 创建年份区间的新向量:
用cut()函数可以快速把年份按10年区间分组,假设你的年份列叫Year:
这里# 定义年份区间和对应标签 df$year_group <- cut(df$Year, breaks = seq(1950, 2020, by = 10), # 按10年间隔设置分界点,可根据数据调整 labels = c("50-59", "60-69", "70-79", "80-89", "90-99", "00-09", "10-19"), include.lowest = TRUE)include.lowest=TRUE是让第一个区间包含起始年份(比如1950)。如果你的年份范围不同,调整breaks里的数值就行。 - 分组分析PctAgreeUS:
用基础R的tapply()可以直接计算每个区间的均值:
用pct_by_group <- tapply(df$PctAgreeUS, df$year_group, mean, na.rm = TRUE)dplyr包会更直观:library(dplyr) group_summary <- df %>% group_by(year_group) %>% summarize(avg_pct = mean(PctAgreeUS, na.rm = TRUE), min_pct = min(PctAgreeUS, na.rm = TRUE), max_pct = max(PctAgreeUS, na.rm = TRUE)) - 可视化变化(可选):
library(ggplot2) ggplot(group_summary, aes(x = year_group, y = avg_pct, group = 1)) + geom_line() + geom_point() + labs(x = "年份区间", y = "PctAgreeUS均值")
Python语言方案
- 加载数据:
import pandas as pd df = pd.read_csv("unvoting.csv") - 创建年份区间列:
用pd.cut()实现分组:# 设置区间和标签 bins = range(1950, 2021, 10) labels = ["50-59", "60-69", "70-79", "80-89", "90-99", "00-09", "10-19"] df["year_group"] = pd.cut(df["Year"], bins=bins, labels=labels, include_lowest=True) - 分组汇总PctAgreeUS:
# 计算每个区间的均值、最值等 group_summary = df.groupby("year_group")["PctAgreeUS"].agg(["mean", "min", "max"]).reset_index() - 可视化变化(可选):
import matplotlib.pyplot as plt plt.plot(group_summary["year_group"], group_summary["mean"], marker='o') plt.xlabel("年份区间") plt.ylabel("PctAgreeUS均值") plt.title("PctAgreeUS随年份区间的变化") plt.show()
内容的提问来源于stack exchange,提问作者Jacob
相关产品推荐
相关产品推荐

