如何在R中将修改后的行更新至已有dataframe?
adaptRisk Function? Hey there! I see the issue with your current adaptRisk function—it's creating a copy of the matching row(s) and modifying that copy instead of updating the original data frame directly. Let's fix that so your changes get integrated back into the original dataset.
First, Let's Recap Your Current Setup
Here's your original function for reference:
adaptRisk <- function(dataframe, sexNum, ageNum, deprivationNum, partnerNum, testResult){ sexRisk = subset(dataframe, sex == sexNum) ageRisk = subset(sexRisk, age == ageNum) depRisk = subset(ageRisk, deprivation == deprivationNum) patientRow = subset(depRisk, partners == partnerNum) if (testResult == "positive") { patientRow$tested <- patientRow$tested + 1 patientRow$infected <- patientRow$infected + 1 } else if (testResult == "negative") { patientRow$tested <- patientRow$tested + 1 } patientRow <- transform(patientRow, risk = infected/tested) return(patientRow) }
Your sample data:
# Data frame head sex age deprivation partners tested infected risk 1 Female 16-19 1-2 0-1 132 1 0.007575758 2 Female 16-19 1-2 2 25 1 0.040000000 3 Female 16-19 1-2 >=3 30 1 0.033333333 4 Female 16-19 3 0-1 80 2 0.025000000 5 Female 16-19 3 2 12 1 0.083333333 6 Female 16-19 3 >=3 18 1 0.055555556 # Dput result structure(list(sex = structure(c(1L, 1L, 1L, 1L, 1L, 1L), .Label = c("Female", "Male"), class = "factor"), age = structure(c(1L, 1L, 1L, 1L, 1L, 1L), .Label = c("16-19", "20-24", "25-34", "35-44"), class = "factor"), deprivation = structure(c(1L, 1L, 1L, 2L, 2L, 2L), .Label = c("1-2", "3", "4-5"), class = "factor"), partners = structure(c(2L, 3L, 1L, 2L, 3L, 1L), .Label = c(">=3", "0-1", "2"), class = "factor"), tested = c(132L, 25L, 30L, 80L, 12L, 18L), infected = c(1L, 1L, 1L, 2L, 1L, 1L), uninfected = c(131L, 24L, 29L, 78L, 11L, 17L), risk = c(0.00757575757575758, 0.04, 0.0333333333333333, 0.025, 0.0833333333333333, 0.0555555555555556)), .Names = c("sex", "age", "deprivation", "partners", "tested", "infected", "uninfected", "risk"), row.names = c(NA, 6L), class = "data.frame")
The Fix: Modify the Original Data Frame Directly
Instead of creating nested subsets (which make copies), we'll first find the index of the matching row in the original data frame, then update those rows directly. Here's the revised function:
adaptRisk <- function(dataframe, sexNum, ageNum, deprivationNum, partnerNum, testResult){ # Find the index of rows matching all your criteria match_rows <- which( dataframe$sex == sexNum & dataframe$age == ageNum & dataframe$deprivation == deprivationNum & dataframe$partners == partnerNum ) # Only proceed if we found matching rows if (length(match_rows) > 0) { # Update the tested count first (always increments) dataframe$tested[match_rows] <- dataframe$tested[match_rows] + 1 # If test is positive, increment infected count too if (testResult == "positive") { dataframe$infected[match_rows] <- dataframe$infected[match_rows] + 1 } # Recalculate risk for the updated rows dataframe$risk[match_rows] <- dataframe$infected[match_rows] / dataframe$tested[match_rows] # Optional: Sync the uninfected column (since it's tested - infected) dataframe$uninfected[match_rows] <- dataframe$tested[match_rows] - dataframe$infected[match_rows] } else { # Warn if no rows match the criteria warning("No rows found matching the provided criteria!") } # Return the full updated data frame return(dataframe) }
How to Use the Revised Function
Since R passes data frames by value, you need to assign the function's output back to your original variable to save the changes:
# Load your data (using the dput result) data <- structure(list(sex = structure(c(1L, 1L, 1L, 1L, 1L, 1L), .Label = c("Female", "Male"), class = "factor"), age = structure(c(1L, 1L, 1L, 1L, 1L, 1L), .Label = c("16-19", "20-24", "25-34", "35-44"), class = "factor"), deprivation = structure(c(1L, 1L, 1L, 2L, 2L, 2L), .Label = c("1-2", "3", "4-5"), class = "factor"), partners = structure(c(2L, 3L, 1L, 2L, 3L, 1L), .Label = c(">=3", "0-1", "2"), class = "factor"), tested = c(132L, 25L, 30L, 80L, 12L, 18L), infected = c(1L, 1L, 1L, 2L, 1L, 1L), uninfected = c(131L, 24L, 29L, 78L, 11L, 17L), risk = c(0.00757575757575758, 0.04, 0.0333333333333333, 0.025, 0.0833333333333333, 0.0555555555555556)), .Names = c("sex", "age", "deprivation", "partners", "tested", "infected", "uninfected", "risk"), row.names = c(NA, 6L), class = "data.frame") # Call the function and update the original data frame data <- adaptRisk(data, "Female", "16-19", "3", "2", "positive") # Check the modified row (row 5) data[5, ]
This will output:
sex age deprivation partners tested infected uninfected risk 5 Female 16-19 3 2 13 2 11 0.1538462
Key Notes
- Index Matching: Using
which()lets us target exactly the rows we need in the original data frame, avoiding copies. - Value Assignment: Remember to reassign the function's output to your original variable (
data <- adaptRisk(...))—otherwise, the changes won't persist. - Error Handling: The
if (length(match_rows) > 0)check prevents errors if no rows match your criteria, and adds a helpful warning. - Data Consistency: The optional
uninfectedupdate ensures all related columns stay in sync.
内容的提问来源于stack exchange,提问作者picador

