正则表达式分组需求:匹配含耗时、距离、ETA的目标字符串
Got it, let's sort out this regex for you! The main issues with your current attempt are likely greedy matching from .+ (which can gobble up more text than you want) and not accounting for all the variations in each segment—like optional hours/minutes, thousand separators in distances, and different unit types.
Here's the working regex that covers all your sample strings and captures the three named groups correctly:
^(?<routearrivalin>(?:\d+\s+h(?:\s+\d+\s+min)?|\d+\s+min))\s+-\s+(?<routedistance>(?:\d{1,3}(?:\.\d{3})*,\d+|\d+)\s+(?:km|m))\s+-\s+Ankunft ca\.\s+(?<routeeta>\d{2}:\d{2})$
Let's break down each part:
routearrivalin: This group handles all three possible time formats in your samples:\d+\s+h\s+\d+\s+min: Matches full time strings like17 h 13 min\d+\s+h: Matches hour-only strings like17 h\d+\s+min: Matches minute-only strings like13 min
The non-capturing groups(?:...)let us group these options without creating extra capture groups.
routedistance: This covers all distance variations:\d{1,3}(?:\.\d{3})*,\d+: Matches numbers with thousand separators and decimals (like1.443,1)\d+: Matches plain integers (like431or3)\s+(?:km|m): Matches the unit (kmorm) with a preceding space
routeeta: This targets the time string specifically by first matching the fixed prefixAnkunft ca., then capturing the time inHH:MMformat with\d{2}:\d{2}.
Testing with your samples:
Let's verify it works for each of your input strings:
- Input:
17 h 13 min - 43,1 km - Ankunft ca. 11:48routearrivalin:17 h 13 minroutedistance:43,1 kmrouteeta:11:48
- Input:
17 h 13 min - 1.443,1 km - Ankunft ca. 11:48routearrivalin:17 h 13 minroutedistance:1.443,1 kmrouteeta:11:48
- Input:
13 min - 431 m - Ankunft ca. 11:48routearrivalin:13 minroutedistance:431 mrouteeta:11:48
- Input:
17 h - 3 km - Ankunft ca. 09:04routearrivalin:17 hroutedistance:3 kmrouteeta:09:04
Quick notes:
- If your input strings might have leading/trailing whitespace, add
\s*at the start (^\s*) and end (\s*$) of the regex. - This assumes all time formats fit the three cases you provided, and distances use
.as thousand separators and,as decimal separators (consistent with your samples).
内容的提问来源于stack exchange,提问作者edba

