如何对列表对象内的所有表执行行求和后列值除以行总和操作
批量计算词频年度占比解决方案
以下提供两种常用数据处理工具的实现方案,均可自动适配列表内不同维度的数据表,无需针对单表单独调整参数:
Python(Pandas)实现
假设存储多表的列表对象名为df_list,实现代码如下:
import pandas as pd def cal_yearly_ratio(df): # 筛选所有非Year的词频列 freq_cols = df.columns.drop("Year") # 计算每行词频总和,处理0值避免除以0报错 row_total = df[freq_cols].sum(axis=1).replace(0, pd.NA) # 词频除以对应行总和得到占比 df[freq_cols] = df[freq_cols].div(row_total, axis=0) # 可按需添加.round(2)指定保留小数位数 return df # 批量处理整个列表 processed_list = [cal_yearly_ratio(df) for df in df_list]
R(tidyverse 生态)实现
假设存储多表的列表对象名为df_list,两种实现方式可选:
借助janitor包简化实现
library(purrr) library(dplyr) library(janitor) processed_list <- df_list |> map(~ .x |> # 排除Year列按行计算占比 adorn_percentages(denominator = "row", na.rm = TRUE, -Year) # 可按需添加 |> mutate(across(-Year, ~round(.x, 2))) 保留指定位数小数 )
纯dplyr实现(无需额外安装包)
library(purrr) library(dplyr) processed_list <- df_list |> map(~ .x |> rowwise() |> mutate(row_total = sum(c_across(-Year), na.rm = TRUE)) |> # 行总和为0时占比自动设为NA,可按需调整替换逻辑 mutate(across(-c(Year, row_total), ~ ifelse(row_total == 0, NA, .x / row_total))) |> select(-row_total) |> ungroup() )
内容的提问来源于stack exchange,提问作者J.Doe
相关产品推荐
相关产品推荐

