You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中从直接与间接URL提取文件扩展名?

问题

需要从以下三种类型的URL中提取文件扩展名,目标是对所有示例URL返回jpg:

  • https://needmode.com/products/350e0f54c3480dc035d6db5e7ef898711d5f4ebc_1683455668.jpg
  • https://dkstatics-public.digikala.com/digikala-products/350e0f54c3480dc035d6db5e7ef898711d5f4ebc_1683455668.jpg?x-oss-process=image/resize,m_lfit,h_800,w_800/quality,q_90
  • https://meghdadit.com/_image.ashx?i=%252ffiles%252fproduct%252f4778c8kbqjb7k18sqydnkztp4yzi0jlaug5j5jtybsmuw0lzq2%255blarge%255d.jpg

当前尝试的Python代码如下:

from urllib.parse import urlparse
import os
img = "IMAGE URL"
parsed_url = urlparse(img)
filename_and_extension = parsed_url.path.rsplit("/", maxsplit=1)[-1]
file_extension = parsed_url.path.rsplit(".", maxsplit=1)[-1].lower()
print("first method: "+file_extension)
filename, file_extension = os.path.splitext(img)
print("second method: "+file_extension)

存在的问题:第一种方法无法处理第三个URL(扩展名在查询参数中),第二种方法无法处理第二个URL(查询参数干扰了扩展名提取)。希望找到一种优先从URL右侧提取扩展名的方案。

解决方案

可以通过从URL末尾反向扫描的方式,精准定位最右侧的合法扩展名,避开路径或查询参数中的干扰字符。具体实现思路:

  1. 反转整个URL,从原URL的末尾位置开始查找第一个.的位置
  2. 从该位置继续扫描,找到第一个?或/(这两个字符会分隔开扩展名和其他URL部分)
  3. 提取中间的字符并反转回来,得到最终的扩展名

对应的Python代码:

def get_file_extension(url):
    reversed_url = url[::-1]
    dot_index = reversed_url.find('.')
    if dot_index == -1:
        return None  # 无扩展名
    
    # 查找扩展名的结束位置(遇到?或/则停止,否则到URL开头)
    end_index = reversed_url.find('?', dot_index)
    if end_index == -1:
        end_index = reversed_url.find('/', dot_index)
    if end_index == -1:
        end_index = len(reversed_url)
    
    # 提取并反转得到扩展名,转小写
    extension = reversed_url[dot_index+1:end_index][::-1].lower()
    # 过滤掉过长/过短的无效扩展名(可选,根据需求调整长度范围)
    if 2 <= len(extension) <= 4:
        return extension
    return None

# 测试示例URL
test_urls = [
    "https://needmode.com/products/350e0f54c3480dc035d6db5e7ef898711d5f4ebc_1683455668.jpg",
    "https://dkstatics-public.digikala.com/digikala-products/350e0f54c3480dc035d6db5e7ef898711d5f4ebc_1683455668.jpg?x-oss-process=image/resize,m_lfit,h_800,w_800/quality,q_90",
    "https://meghdadit.com/_image.ashx?i=%252ffiles%252fproduct%252f4778c8kbqjb7k18sqydnkztp4yzi0jlaug5j5jtybsmuw0lzq2%255blarge%255d.jpg"
]

for url in test_urls:
    print(f"URL: {url}")
    print(f"提取的扩展名: {get_file_extension(url)}\n")

这段代码对三个示例URL都会返回jpg,同时能兼容扩展名出现在路径、查询参数等不同位置的情况。

内容的提问来源于stack exchange,提问作者Martin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.30 11:27:04