在Spark RDD中使用正则表达式提取字符串时间戳的问题
Hey there! Let's get that 20:33:05 time extracted from your log properly. Since you mentioned your initial attempts didn't work, let's cover two common regex approaches—one that strictly validates the time format, and a simpler version for when you know your log's time structure is consistent.
Option 1: Strictly Validated Regex (Matches Valid HH:MM:SS)
This regex ensures you only capture valid 24-hour time values (hours 00-23, minutes/seconds 00-59):
\b(?:[01]\d|2[0-3]):(?:[0-5]\d):(?:[0-5]\d)\b
Breakdown of the regex:
\b: Word boundary, ensures we don't pick up partial time strings (like123:45:67ora20:33:05b)(?:[01]\d|2[0-3]): Non-capturing group for hours:[01]\d: Matches 00-192[0-3]: Matches 20-23
(?:[0-5]\d): Non-capturing group for minutes and seconds, matches 00-59- The colons
:are literal characters matching the time separators
Option 2: Simplified Regex (For Consistent Log Formats)
If you're certain your log always uses the HH:MM:SS pattern (even if you don't need to validate validity), this simpler regex works:
\b\d{2}:\d{2}:\d{2}\b
Breakdown:
\d{2}: Matches exactly two digits (for hours, minutes, seconds):: Literal colon separators\b: Word boundary to avoid partial matches
Example Usage
Suppose your log line looks like this:
[2024-05-20 20:33:05] INFO: User completed action
Using either regex above will correctly extract 20:33:05. If you're using Python, here's a quick implementation:
import re log_line = "[2024-05-20 20:33:05] INFO: User completed action" time_regex = r"\b(?:[01]\d|2[0-3]):(?:[0-5]\d):(?:[0-5]\d)\b" match = re.search(time_regex, log_line) if match: print("Extracted time:", match.group()) # Output: Extracted time: 20:33:05
If your log has a more unique structure (like time prefixed with extra characters), just adjust the regex to account for those—feel free to share more details if you need further tweaks!
内容的提问来源于stack exchange,提问作者Keshav Pradeep Ramanath

