You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python从指定URL模式列表提取页面ID?求适用工具库

问题描述

我有一组包含网页Page ID的URL模式列表:

patterns = [
    "https://www.my-website.com/articles/$1/show",
    "https://www.my-website.com/groups/$1/articles/$2",
    "https://www.my-website.com/blogs/$1/show",
    "https://www.my-website.com/groups/$1/documents/$2"
]

其中$1、$2等为对应位置的Page ID。我需要根据输入的WebURL字符串提取对应的Page ID,示例如下:

input = "https://www.my-website.com/articles/asojd1277/show"
output = "asojd1277"

input = "https://www.my-website.com/groups/hsaid123/articles/hqwoj8239"
output = "hqwoj8239"

我尝试过re.match(正则包)但不符合需求,请问是否有合适的Python库可实现该功能?


解决方案

推荐两个实用的Python库来高效处理这类URL匹配与参数提取:

1. Werkzeug(专业路由匹配)

Werkzeug是一款WSGI工具库,自带的路由系统可以精准定义URL规则并提取参数,完全适配你的需求。

使用步骤:

  • 安装:pip install werkzeug
  • 代码示例:
from werkzeug.routing import Map, Rule

# 把原模式中的$1、$2替换为命名占位符,定义路由规则
url_map = Map([
    Rule('/articles/<page_id>/show', endpoint='article_show'),
    Rule('/groups/<group_id>/articles/<page_id>', endpoint='group_article'),
    Rule('/blogs/<page_id>/show', endpoint='blog_show'),
    Rule('/groups/<group_id>/documents/<page_id>', endpoint='group_document'),
])

def extract_page_id(input_url):
    adapter = url_map.bind('www.my-website.com')
    try:
        _, values = adapter.match(input_url.split('://')[1])
        # 返回目标Page ID(对应示例中最后一个参数的需求)
        return values.get('page_id')
    except:
        return None

# 测试示例
print(extract_page_id("https://www.my-website.com/articles/asojd1277/show"))  # 输出: asojd1277
print(extract_page_id("https://www.my-website.com/groups/hsaid123/articles/hqwoj8239"))  # 输出: hqwoj8239

2. parse(轻量字符串匹配)

parse库的语法和你现有的$1、$2模式高度契合,上手简单,适合快速实现需求。

使用步骤:

  • 安装:pip install parse
  • 代码示例:
from parse import parse

# 将原模式中的$1替换为{},按顺序匹配参数
patterns = [
    "https://www.my-website.com/articles/{}/show",
    "https://www.my-website.com/groups/{}/articles/{}",
    "https://www.my-website.com/blogs/{}/show",
    "https://www.my-website.com/groups/{}/documents/{}"
]

def extract_page_id(input_url):
    for pattern in patterns:
        result = parse(pattern, input_url)
        if result:
            # 返回最后一个匹配到的参数(对应示例需求)
            return result[-1]
    return None

# 测试示例
print(extract_page_id("https://www.my-website.com/articles/asojd1277/show"))  # 输出: asojd1277
print(extract_page_id("https://www.my-website.com/groups/hsaid123/articles/hqwoj8239"))  # 输出: hqwoj8239

补充:正则的可行写法

如果之前用re.match未成功,大概率是规则编写问题。针对你的需求,正则也可以实现,但多模式场景下维护成本较高。示例:

import re

def extract_page_id(input_url):
    # 针对每个模式编写正则
    regex_patterns = [
        r"https://www.my-website.com/articles/(\w+)/show",
        r"https://www.my-website.com/groups/\w+/articles/(\w+)",
        r"https://www.my-website.com/blogs/(\w+)/show",
        r"https://www.my-website.com/groups/\w+/documents/(\w+)"
    ]
    for pattern in regex_patterns:
        match = re.match(pattern, input_url)
        if match:
            return match.group(1)
    return None

内容的提问来源于stack exchange,提问作者Vishnukk

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.27 00:34:55