You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python提取URL指定片段并保留原始斜杠格式的方法

Extracting Exact Slash-Formatted Segment from URL

Got it, let's fix this for you! The problem with using split('/')[3:5] is twofold: you're only grabbing two of the three target segments, and splitting removes the slashes you want to preserve. Here are a couple of reliable ways to get /1234/44/222 exactly as you need it:

Method 1: Fix the Split & Join Approach

First, let's correct the split indexing and reconstruct the slash-formatted string properly:

url = "https://espn.com/1234/44/222/mlb/standings"
split_parts = url.split('/')
# The segments we want are at indices 3, 4, 5: '1234', '44', '222'
target_segments = split_parts[3:6]
# Reconstruct with leading slash and internal slashes
result = '/' + '/'.join(target_segments)
print(result)  # Output: /1234/44/222

This works because we first split the URL into parts, grab the three numeric segments, then join them with slashes and prepend the leading slash to match your desired format.

Method 2: Use Regular Expressions (More Flexible)

If you want a solution that adapts better to potential URL structure changes (like different domain names), regex is a great option. We can match the domain followed by the three slash-separated segments:

import re

url = "https://espn.com/1234/44/222/mlb/standings"
# The regex matches the domain part, then captures the three slash-separated segments
match = re.search(r'//[^/]+(/\d+/\d+/\d+)', url)
if match:
    result = match.group(1)
    print(result)  # Output: /1234/44/222

The regex //[^/]+ targets the domain after //, then (/\d+/\d+/\d+) captures exactly the three numeric segments with their slashes. If your segments aren't always numbers, replace \d+ with [^/]+ to match any characters between slashes.

Method 3: String Slicing (For Fixed URL Structure)

If you know the URL structure will always stay the same, you can use string slicing to directly extract the desired part:

url = "https://espn.com/1234/44/222/mlb/standings"
# Find the end of the domain part
domain_end = url.find('//espn.com') + len('//espn.com')
# Find the position of the fourth slash after the domain (to mark the end of our target segment)
slash_count = 0
target_end = domain_end
for idx in range(domain_end, len(url)):
    if url[idx] == '/':
        slash_count += 1
        if slash_count == 4:
            target_end = idx
            break
result = url[domain_end:target_end]
print(result)  # Output: /1234/44/222

This is more rigid, but works perfectly if you're certain the URL won't change its structure.

内容的提问来源于stack exchange,提问作者skimchi1993

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 07:04:45