从指定URL列表中提取域名与URL后缀字符串
提取URL域名与结尾后缀的Python实现
我们可以利用Python标准库中的urllib.parse模块解析域名,再通过字符串处理提取结尾后缀。以下是针对给定URL列表的实现代码:
from urllib.parse import urlparse # 待处理的URL列表 url_list = ['https://blog.hubspot.com/marketing/parts-url', 'https://www.almabetter.com/enrollments', 'https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.DataFrame.rename.html', 'https://www.programiz.com/python-programming/list'] for url in url_list: # 解析URL获取域名 parsed_url = urlparse(url) domain = parsed_url.netloc # 提取结尾字符串(后缀) path = parsed_url.path last_segment = path.split('/')[-1] suffix = last_segment.split('.')[-1] if '.' in last_segment else '' print(f"URL: {url}") print(f"域名: {domain}") print(f"结尾字符串: {suffix}\n")
代码说明
- 域名提取:通过
urlparse()解析URL后,netloc属性直接返回域名部分。 - 结尾后缀提取:先获取URL的路径部分,分割出最后一段内容;如果该内容包含点号,则取点号后的最后一段作为后缀,否则返回空字符串。
运行结果
URL: https://blog.hubspot.com/marketing/parts-url 域名: blog.hubspot.com 结尾字符串: URL: https://www.almabetter.com/enrollments 域名: www.almabetter.com 结尾字符串: URL: https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.DataFrame.rename.html 域名: pandas.pydata.org 结尾字符串: html URL: https://www.programiz.com/python-programming/list 域名: www.programiz.com 结尾字符串:
内容的提问来源于stack exchange,提问作者Syed Sajjad Askari
相关产品推荐
相关产品推荐

