R语言处理字符串向量前导零:替换规则与特殊格式兼容(Tidyverse优先)
Got it, let's tackle this step by step using tidyverse's stringr package—perfect for this kind of string manipulation. Your main challenges are handling two cases of leading zeros, fixing the missing space before AM/PM in the last entry, and making sure we don't mess up minutes or valid two-digit hours like 10 or 12.
Step 1: Load Tidyverse and Define Your Input
First, let's set up our environment and input vector:
library(tidyverse) # Your original character vector time_vec <- c("TIME 00:35 AM 08:34 AM 10:10 AM 10:00 PM 09:45 PM 02:00 PM 07:47 PM 12:22AM")
Step 2: Fix Missing Space Before AM/PM
First, let's standardize all time entries to have a space before AM/PM. The last entry (12:22AM) is missing this, so we'll use a regex to find cases where digits are directly followed by AM/PM and insert a space:
# Add space between time and AM/PM where it's missing time_with_spaces <- str_replace_all(time_vec, "(\\d{2}:\\d{2})(AM|PM)", "\\1 \\2")
This regex captures the time part (\\d{2}:\\d{2}) and the AM/PM part separately, then puts them back together with a space in between.
Step 3: Handle Leading Zeros in Hours
Now we'll fix the leading zeros with two targeted replacements:
- Replace
00(which should become 12) - Remove single leading zeros (like
08→8,02→2)
We'll use str_replace_all with regex lookaheads to make sure we only target hour values (not minutes like 10:00):
final_formatted_time <- time_with_spaces %>% # Replace "00:" with "12:" (only when it's an hour) str_replace_all("\\b00(?=:)", "12") %>% # Remove single leading zero from hours (e.g., 08: → 8:) str_replace_all("\\b0(?=[1-9]:)", "")
Let's break down the regex here:
\\b00(?=:):\\bensures we're matching a whole word (so we don't accidentally target part of another number), and(?=:)is a positive lookahead that checks for a colon immediately after—this guarantees we're only replacing the hour value00.\\b0(?=[1-9]:): Similarly, this matches a single0followed by a 1-9 digit and a colon, which is exactly the single leading zero we want to remove from hours like 08 or 09.
Check the Result
If you run final_formatted_time, you'll get exactly what you wanted:
final_formatted_time # Output: "TIME 12:35 AM 8:34 AM 10:10 AM 10:00 PM 9:45 PM 2:00 PM 7:47 PM 12:22 AM"
Why Your Initial str_remove Didn't Work
str_remove is great for deleting specific patterns, but it can't differentiate between the two leading zero cases we have here (replacing 00 with 12 vs. removing a single 0). We need conditional replacement, which str_replace_all handles perfectly with targeted regex patterns.
内容的提问来源于stack exchange,提问作者Johnny Thomas

