如何在Stata中按年份分组计算变量总和并生成汇总变量?
按年份计算变量总和并绘制直方图
问题说明
我的数据集包含48个变量:前12个对应2000年,第13-24个对应2001年,第25-36个对应2002年,第37-48个对应2003年。示例数据如下:
| ID | var1 | var2 | var3 | ... | var12 | ... | var48 |
|---|---|---|---|---|---|---|---|
| xx | 0 | 0 | 1 | ... | 1 | ... | 0 |
| yy | 1 | 0 | 0 | ... | 9 | ... | 0 |
| zz | 3 | 2 | 1 | ... | 0 | ... | 0 |
需要将每年对应变量的总和分别存入tot_2000、tot_2001、tot_2002、tot_2003四个变量,用于绘制直方图。其中tot_2000的示例输出如下:
| tot_2000 |
|---|
| 18 |
实现方案
1. 使用R语言
# 读取数据集(假设数据框名为df) # df <- read.csv("your_data.csv") # 计算各年份总和 df$tot_2000 <- rowSums(df[, 1:12]) # 前12列对应2000年 df$tot_2001 <- rowSums(df[, 13:24]) # 第13-24列对应2001年 df$tot_2002 <- rowSums(df[, 25:36]) # 第25-36列对应2002年 df$tot_2003 <- rowSums(df[, 37:48]) # 第37-48列对应2003年 # 绘制单一年份直方图(以2000年为例) hist(df$tot_2000, main = "2000年变量总和直方图", xlab = "总和", col = "lightblue") # 批量绘制四年直方图(2x2布局) par(mfrow = c(2, 2)) hist(df$tot_2000, main = "2000年总和", xlab = "数值", col = "lightblue") hist(df$tot_2001, main = "2001年总和", xlab = "数值", col = "lightgreen") hist(df$tot_2002, main = "2002年总和", xlab = "数值", col = "lightyellow") hist(df$tot_2003, main = "2003年总和", xlab = "数值", col = "lightpink") par(mfrow = c(1, 1)) # 恢复默认布局
2. 使用Python(Pandas + Matplotlib)
import pandas as pd import matplotlib.pyplot as plt # 读取数据集(假设数据文件名为your_data.csv) # df = pd.read_csv("your_data.csv") # 计算各年份总和(注意Pandas索引从0开始) df['tot_2000'] = df.iloc[:, 0:12].sum(axis=1) df['tot_2001'] = df.iloc[:, 12:24].sum(axis=1) df['tot_2002'] = df.iloc[:, 24:36].sum(axis=1) df['tot_2003'] = df.iloc[:, 36:48].sum(axis=1) # 绘制单一年份直方图(以2000年为例) plt.figure(figsize=(6,4)) plt.hist(df['tot_2000'], bins=10, color='lightblue', edgecolor='black') plt.title('2000年变量总和直方图') plt.xlabel('总和') plt.ylabel('频数') plt.show() # 批量绘制四年直方图(2x2布局) fig, axes = plt.subplots(2, 2, figsize=(10,8)) years = ['tot_2000', 'tot_2001', 'tot_2002', 'tot_2003'] colors = ['lightblue', 'lightgreen', 'lightyellow', 'lightpink'] titles = ['2000年总和', '2001年总和', '2002年总和', '2003年总和'] for ax, year, color, title in zip(axes.flat, years, colors, titles): ax.hist(df[year], bins=10, color=color, edgecolor='black') ax.set_title(title) ax.set_xlabel('数值') ax.set_ylabel('频数') plt.tight_layout() plt.show()
3. 使用Stata
* 读取数据集(假设数据文件名为your_data.dta) * use "your_data.dta", clear * 计算各年份总和 gen tot_2000 = rowtotal(var1-var12) // 前12个变量对应2000年 gen tot_2001 = rowtotal(var13-var24) // 第13-24个变量对应2001年 gen tot_2002 = rowtotal(var25-var36) // 第25-36个变量对应2002年 gen tot_2003 = rowtotal(var37-var48) // 第37-48个变量对应2003年 * 绘制单一年份直方图(以2000年为例) histogram tot_2000, title("2000年变量总和直方图") xlabel("总和") color(lightblue) * 批量绘制四年直方图 foreach year in tot_2000 tot_2001 tot_2002 tot_2003 { histogram `year', title("`year'直方图") xlabel("数值") color(lightblue) }
内容的提问来源于stack exchange,提问作者tryingtogetsmth
相关产品推荐
相关产品推荐

