Python提取URL指定片段并保留原始斜杠格式的方法
Got it, let's fix this for you! The problem with using split('/')[3:5] is twofold: you're only grabbing two of the three target segments, and splitting removes the slashes you want to preserve. Here are a couple of reliable ways to get /1234/44/222 exactly as you need it:
Method 1: Fix the Split & Join Approach
First, let's correct the split indexing and reconstruct the slash-formatted string properly:
url = "https://espn.com/1234/44/222/mlb/standings" split_parts = url.split('/') # The segments we want are at indices 3, 4, 5: '1234', '44', '222' target_segments = split_parts[3:6] # Reconstruct with leading slash and internal slashes result = '/' + '/'.join(target_segments) print(result) # Output: /1234/44/222
This works because we first split the URL into parts, grab the three numeric segments, then join them with slashes and prepend the leading slash to match your desired format.
Method 2: Use Regular Expressions (More Flexible)
If you want a solution that adapts better to potential URL structure changes (like different domain names), regex is a great option. We can match the domain followed by the three slash-separated segments:
import re url = "https://espn.com/1234/44/222/mlb/standings" # The regex matches the domain part, then captures the three slash-separated segments match = re.search(r'//[^/]+(/\d+/\d+/\d+)', url) if match: result = match.group(1) print(result) # Output: /1234/44/222
The regex //[^/]+ targets the domain after //, then (/\d+/\d+/\d+) captures exactly the three numeric segments with their slashes. If your segments aren't always numbers, replace \d+ with [^/]+ to match any characters between slashes.
Method 3: String Slicing (For Fixed URL Structure)
If you know the URL structure will always stay the same, you can use string slicing to directly extract the desired part:
url = "https://espn.com/1234/44/222/mlb/standings" # Find the end of the domain part domain_end = url.find('//espn.com') + len('//espn.com') # Find the position of the fourth slash after the domain (to mark the end of our target segment) slash_count = 0 target_end = domain_end for idx in range(domain_end, len(url)): if url[idx] == '/': slash_count += 1 if slash_count == 4: target_end = idx break result = url[domain_end:target_end] print(result) # Output: /1234/44/222
This is more rigid, but works perfectly if you're certain the URL won't change its structure.
内容的提问来源于stack exchange,提问作者skimchi1993

