You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python下载带签名与过期时间的S3链接文件?

Hey there, let's troubleshoot why that signed S3 link works in your browser but throws a 403 error when you use urllib.request.urlretrieve. This is a common issue, and it usually boils down to differences in how browsers vs. Python's HTTP libraries handle the request. Here are the most likely fixes:

1. URL Encoding Mismatches

Browsers automatically handle decoding/encoding for special characters in the Signature parameter (like %2B or %3D), but urllib might re-encode parts of the URL, breaking S3's signature validation.

A quick, reliable fix is to use the requests library instead—it handles URL edge cases much more gracefully:

import requests

signed_url = "http://s3.amazonaws.com/your-bucket/path/to/file?AWSAccessKeyId=...&Expires=...&Signature=..."

# Stream the download to avoid loading large files into memory
response = requests.get(signed_url, stream=True)
response.raise_for_status()  # Catch HTTP errors early

with open("downloaded_file.ext", "wb") as f:
    for chunk in response.iter_content(chunk_size=8192):
        f.write(chunk)

2. Missing Request Headers

S3's signature validation checks more than just the URL—sometimes it verifies request headers too. Browsers send a User-Agent header by default, but urllib uses a minimal default that might be rejected by S3.

If you want to stick with urllib, add a browser-like User-Agent to your request:

import urllib.request

signed_url = "your-signed-s3-link-here"

# Mimic a standard browser's request headers
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36"
}

# Create a custom request with headers
req = urllib.request.Request(signed_url, headers=headers)

# Execute the request and save the file
with urllib.request.urlopen(req) as response, open("downloaded_file.ext", "wb") as out_file:
    out_file.write(response.read())

3. Issues with Manual Boto3 Parsing

If you tried using boto3 by manually extracting bucket names, keys, and signature parameters, you might have messed up the URL parsing (S3 URLs come in two formats: bucket.s3.amazonaws.com/path or s3.amazonaws.com/bucket/path).

That said, you don't need boto3 to download from a pre-signed link—the link already contains all required authentication. Using a simple HTTP library like requests is far simpler. If you still want to use boto3, here's how to correctly parse the URL:

from urllib.parse import urlparse, parse_qs
import boto3

signed_url = "your-signed-s3-link-here"
parsed_url = urlparse(signed_url)
query_params = parse_qs(parsed_url.query)

# Extract bucket name based on URL structure
if parsed_url.netloc == "s3.amazonaws.com":
    # Format: s3.amazonaws.com/bucket/path/to/file
    bucket_name = parsed_url.path.lstrip("/").split("/")[0]
    key = "/".join(parsed_url.path.lstrip("/").split("/")[1:])
else:
    # Format: bucket.s3.amazonaws.com/path/to/file
    bucket_name = parsed_url.netloc.split(".")[0]
    key = parsed_url.path.lstrip("/")

# Use boto3 to fetch the object (overkill, but works)
s3 = boto3.client("s3")
try:
    response = s3.get_object(Bucket=bucket_name, Key=key)
    with open("downloaded_file.ext", "wb") as f:
        f.write(response["Body"].read())
except Exception as e:
    print(f"Error with boto3: {e}")

Final Notes

The most reliable fix is using requests—it handles encoding and headers automatically, eliminating the 403 error. If you can't install requests, the custom urllib request with a proper User-Agent should work too.

内容的提问来源于stack exchange,提问作者user3776598

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 08:15:32