如何在Python中从直接与间接URL提取文件扩展名?
问题
需要从以下三种类型的URL中提取文件扩展名,目标是对所有示例URL返回jpg:
https://needmode.com/products/350e0f54c3480dc035d6db5e7ef898711d5f4ebc_1683455668.jpghttps://dkstatics-public.digikala.com/digikala-products/350e0f54c3480dc035d6db5e7ef898711d5f4ebc_1683455668.jpg?x-oss-process=image/resize,m_lfit,h_800,w_800/quality,q_90https://meghdadit.com/_image.ashx?i=%252ffiles%252fproduct%252f4778c8kbqjb7k18sqydnkztp4yzi0jlaug5j5jtybsmuw0lzq2%255blarge%255d.jpg
当前尝试的Python代码如下:
from urllib.parse import urlparse import os img = "IMAGE URL" parsed_url = urlparse(img) filename_and_extension = parsed_url.path.rsplit("/", maxsplit=1)[-1] file_extension = parsed_url.path.rsplit(".", maxsplit=1)[-1].lower() print("first method: "+file_extension) filename, file_extension = os.path.splitext(img) print("second method: "+file_extension)
存在的问题:第一种方法无法处理第三个URL(扩展名在查询参数中),第二种方法无法处理第二个URL(查询参数干扰了扩展名提取)。希望找到一种优先从URL右侧提取扩展名的方案。
解决方案
可以通过从URL末尾反向扫描的方式,精准定位最右侧的合法扩展名,避开路径或查询参数中的干扰字符。具体实现思路:
- 反转整个URL,从原URL的末尾位置开始查找第一个
.的位置 - 从该位置继续扫描,找到第一个
?或/(这两个字符会分隔开扩展名和其他URL部分) - 提取中间的字符并反转回来,得到最终的扩展名
对应的Python代码:
def get_file_extension(url): reversed_url = url[::-1] dot_index = reversed_url.find('.') if dot_index == -1: return None # 无扩展名 # 查找扩展名的结束位置(遇到?或/则停止,否则到URL开头) end_index = reversed_url.find('?', dot_index) if end_index == -1: end_index = reversed_url.find('/', dot_index) if end_index == -1: end_index = len(reversed_url) # 提取并反转得到扩展名,转小写 extension = reversed_url[dot_index+1:end_index][::-1].lower() # 过滤掉过长/过短的无效扩展名(可选,根据需求调整长度范围) if 2 <= len(extension) <= 4: return extension return None # 测试示例URL test_urls = [ "https://needmode.com/products/350e0f54c3480dc035d6db5e7ef898711d5f4ebc_1683455668.jpg", "https://dkstatics-public.digikala.com/digikala-products/350e0f54c3480dc035d6db5e7ef898711d5f4ebc_1683455668.jpg?x-oss-process=image/resize,m_lfit,h_800,w_800/quality,q_90", "https://meghdadit.com/_image.ashx?i=%252ffiles%252fproduct%252f4778c8kbqjb7k18sqydnkztp4yzi0jlaug5j5jtybsmuw0lzq2%255blarge%255d.jpg" ] for url in test_urls: print(f"URL: {url}") print(f"提取的扩展名: {get_file_extension(url)}\n")
这段代码对三个示例URL都会返回jpg,同时能兼容扩展名出现在路径、查询参数等不同位置的情况。
内容的提问来源于stack exchange,提问作者Martin
相关产品推荐
相关产品推荐

