如何编写Python URL对比函数实现忽略路径后缀识别页面类型
URL路径匹配需求实现方案
你可以使用Python字符串内置的startswith()方法实现前缀匹配,同时为了避免路径误匹配(比如把/questionsabc这种不符合规则的路径识别为question类型),可以给基准路径统一增加末尾斜杠做边界限制。
修改后的代码如下:
def bbox_page_type(url): # 定义基准路径,末尾统一加/做边界限制 home = 'https://forum.bouyguestelecom.fr/' questions_list = 'https://forum.bouyguestelecom.fr/categories/' question = 'https://forum.bouyguestelecom.fr/questions/' # 首页要求完全匹配 if url == home: print("home.") # 匹配分类列表路径,兼容基准路径本身(不带末尾/)和带后续路径的情况 elif url.startswith(questions_list) or url == questions_list.rstrip('/'): print("question list.") # 匹配问题详情路径 elif url.startswith(question) or url == question.rstrip('/'): print("question.") # 其余均为未知类型 else: print("UNKNOWN.") # 测试示例 if __name__ == "__main__": # 带后续路径的分类列表页 bbox_page_type('https://forum.bouyguestelecom.fr/categories/手机套餐') # 带后续路径的问题详情页 bbox_page_type('https://forum.bouyguestelecom.fr/questions/123456-如何查话费') # 无后续路径的分类列表根路径 bbox_page_type('https://forum.bouyguestelecom.fr/categories') # 首页 bbox_page_type('https://forum.bouyguestelecom.fr/') # 未知路径 bbox_page_type('https://forum.bouyguestelecom.fr/user/123')
代码说明
- 替换原有的全等判断为前缀匹配,只要URL以对应基准路径开头就会命中规则,自动忽略基准路径后接的所有内容,包括子路径、查询参数等
- 用
if-elif分支结构替代原有的多个独立if,匹配到对应类型后直接终止判断,减少无效运算 - 额外增加了无末尾斜杠的基准路径兼容逻辑,保证根路径本身也能正确命中规则
内容的提问来源于stack exchange,提问作者ossama assaghir
相关产品推荐
相关产品推荐

