Chrome可访问的TIF文件无法用wget下载,报401未授权
Hey there! Let's figure out why Chrome can download those TIF files without login issues, but wget is throwing a 401 error. The key here is that browsers send extra headers and sometimes session cookies that the server expects—wget doesn't do this by default. Here's how to fix it:
1. Capture Chrome's Request Details
First, let's see exactly what Chrome sends to the server when it downloads the file. This will give us the headers/cookies we need to replicate in wget:
- Open Chrome and navigate to your TIF file URL.
- Press
F12to open DevTools, then switch to the Network tab. - Find the entry for your TIF file (it should be near the top), right-click it, and select Copy > Copy as cURL (bash).
2. Convert cURL Command to Wget
The copied cURL command includes all the headers and cookies Chrome used. We can translate this to wget parameters:
- Extract the
User-Agent,Accept, andCookieheaders from the cURL command. - Use wget's
--headerflag to add each one. Here's an example (replace the placeholder values with what you copied):wget --no-check-certificate \ --header="User-Agent: Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" \ --header="Accept: image/webp,image/apng,image/svg+xml,image/*,*/*;q=0.8" \ --header="Cookie: paste_your_cookie_string_here" \ https://ulib.aub.edu.lb/nahar/images2/7810W2/78101001.TIF
3. Test Simplified Versions First
You might not need all headers—start with the simplest fix to save time:
- Try just the User-Agent first: Some servers block requests with non-browser User-Agents (wget uses a default one like
Wget/1.21.3). Run this:wget --no-check-certificate --user-agent="Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" https://ulib.aub.edu.lb/nahar/images2/7810W2/78101001.TIF - If that fails, add the
Cookieheader from your cURL copy. The server might be setting a session cookie even without login (for tracking or access control).
4. Alternative: Use cURL Directly
If converting to wget feels tricky, you can use the copied cURL command directly (just add -O to save the file with its original name):
curl --insecure -O https://ulib.aub.edu.lb/nahar/images2/7810W2/78101001.TIF \ -H "User-Agent: Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" \ -H "Accept: image/webp,image/apng,image/svg+xml,image/*,*/*;q=0.8" \ -H "Cookie: paste_your_cookie_string_here"
(--insecure does the same as wget's --no-check-certificate to bypass SSL certificate checks.)
Let me know if any of these steps work for you—happy to help tweak things further if needed!
内容的提问来源于stack exchange,提问作者Abuflaan

