如何用Python从指定URL模式列表提取页面ID?求适用工具库
问题描述
我有一组包含网页Page ID的URL模式列表:
patterns = [ "https://www.my-website.com/articles/$1/show", "https://www.my-website.com/groups/$1/articles/$2", "https://www.my-website.com/blogs/$1/show", "https://www.my-website.com/groups/$1/documents/$2" ]
其中$1、$2等为对应位置的Page ID。我需要根据输入的WebURL字符串提取对应的Page ID,示例如下:
input = "https://www.my-website.com/articles/asojd1277/show" output = "asojd1277" input = "https://www.my-website.com/groups/hsaid123/articles/hqwoj8239" output = "hqwoj8239"
我尝试过re.match(正则包)但不符合需求,请问是否有合适的Python库可实现该功能?
解决方案
推荐两个实用的Python库来高效处理这类URL匹配与参数提取:
1. Werkzeug(专业路由匹配)
Werkzeug是一款WSGI工具库,自带的路由系统可以精准定义URL规则并提取参数,完全适配你的需求。
使用步骤:
- 安装:
pip install werkzeug - 代码示例:
from werkzeug.routing import Map, Rule # 把原模式中的$1、$2替换为命名占位符,定义路由规则 url_map = Map([ Rule('/articles/<page_id>/show', endpoint='article_show'), Rule('/groups/<group_id>/articles/<page_id>', endpoint='group_article'), Rule('/blogs/<page_id>/show', endpoint='blog_show'), Rule('/groups/<group_id>/documents/<page_id>', endpoint='group_document'), ]) def extract_page_id(input_url): adapter = url_map.bind('www.my-website.com') try: _, values = adapter.match(input_url.split('://')[1]) # 返回目标Page ID(对应示例中最后一个参数的需求) return values.get('page_id') except: return None # 测试示例 print(extract_page_id("https://www.my-website.com/articles/asojd1277/show")) # 输出: asojd1277 print(extract_page_id("https://www.my-website.com/groups/hsaid123/articles/hqwoj8239")) # 输出: hqwoj8239
2. parse(轻量字符串匹配)
parse库的语法和你现有的$1、$2模式高度契合,上手简单,适合快速实现需求。
使用步骤:
- 安装:
pip install parse - 代码示例:
from parse import parse # 将原模式中的$1替换为{},按顺序匹配参数 patterns = [ "https://www.my-website.com/articles/{}/show", "https://www.my-website.com/groups/{}/articles/{}", "https://www.my-website.com/blogs/{}/show", "https://www.my-website.com/groups/{}/documents/{}" ] def extract_page_id(input_url): for pattern in patterns: result = parse(pattern, input_url) if result: # 返回最后一个匹配到的参数(对应示例需求) return result[-1] return None # 测试示例 print(extract_page_id("https://www.my-website.com/articles/asojd1277/show")) # 输出: asojd1277 print(extract_page_id("https://www.my-website.com/groups/hsaid123/articles/hqwoj8239")) # 输出: hqwoj8239
补充:正则的可行写法
如果之前用re.match未成功,大概率是规则编写问题。针对你的需求,正则也可以实现,但多模式场景下维护成本较高。示例:
import re def extract_page_id(input_url): # 针对每个模式编写正则 regex_patterns = [ r"https://www.my-website.com/articles/(\w+)/show", r"https://www.my-website.com/groups/\w+/articles/(\w+)", r"https://www.my-website.com/blogs/(\w+)/show", r"https://www.my-website.com/groups/\w+/documents/(\w+)" ] for pattern in regex_patterns: match = re.match(pattern, input_url) if match: return match.group(1) return None
内容的提问来源于stack exchange,提问作者Vishnukk
相关产品推荐
相关产品推荐

