如何在R语言中自动识别并修正数据集多行列数值倍数以匹配对应列小数格式
me Column Values Based on Event in R Great question! Since each row's me value needs a specific multiplier to align its decimal format with the act column, we can set up an automated way to map each Event to its required multiplier in R. Below are two straightforward, scalable methods to solve this:
Method 1: Using a Multiplier Mapping Table (Best for Large/Growing Datasets)
This approach is perfect if your dataset will expand with more events later—you only need to update the mapping table instead of rewriting core logic.
# Install and load dplyr if you haven't already # install.packages("dplyr") library(dplyr) # Create a table that links each Event to its needed multiplier multiplier_map <- tibble( Event = c( "Labour Costs YoY", "Unemployment Change (000's)", "Unemployment Rate", "Jobseekers Net Change" ), multiplier = c(100, 10, 100, 1) ) # Join the multiplier map to your dataset and adjust the `me` column ds2_updated <- ds2 %>% left_join(multiplier_map, by = "Event") %>% mutate(me = me * multiplier) %>% select(-multiplier) # Remove the helper column if not needed # View the corrected dataset ds2_updated
Method 2: Using case_when (Simple for Smaller Event Lists)
If you only have a handful of events to handle, you can apply the multipliers directly with case_when in a mutate call:
library(dplyr) ds2_updated <- ds2 %>% mutate(me = case_when( Event == "Labour Costs YoY" ~ me * 100, Event == "Unemployment Change (000's)" ~ me * 10, Event == "Unemployment Rate" ~ me * 100, Event == "Jobseekers Net Change" ~ me * 1, TRUE ~ me # Leave `me` unchanged for any unexpected events )) # Check the result ds2_updated
Result Verification
Both methods will produce the corrected dataset where me matches the decimal format of act:
# A tibble: 4 × 3 Event act me <fct> <dbl> <dbl> 1 Labour Costs YoY 2.33 2.64 2 Unemployment Change (000's) -5.17 -4.49 3 Unemployment Rate 8.86 9.00 4 Jobseekers Net Change 11.3 9.57
Key Benefit
Both approaches are fully automated—no manual row-by-row edits required. As your dataset grows, just update the multiplier map (Method 1) or add new conditions to case_when (Method 2), and the code will handle the rest.
内容的提问来源于stack exchange,提问作者Jurgen

