如何通过命令行用凭据与Cookie访问私有网站(含wget配置)
Hey there! Since you've got 2000 links to a private database site that requires cookies for access, let's walk through how to configure wget (and curl, as a solid alternative) to handle this properly. Here's what you need to know:
Wget works best with Netscape-format cookie files (the old-style plain text format, not JSON). You have two easy ways to get this:
Option A: Export from Your Browser
If you're already logged into the private site in your browser:
- For Chrome/Firefox: Use a browser extension like "Export Cookies" to save your cookies in Netscape format, or go to Developer Tools (F12) → Application/Storage tab → Cookies, then copy relevant entries into a text file following this structure:
Make sure the file starts with the# Netscape HTTP Cookie File your-private-site.com TRUE / FALSE 1718923200 session_id abc123xyz your-private-site.com TRUE / FALSE 1718923200 user_token 987654321# Netscape HTTP Cookie Fileheader—wget looks for this to recognize the format.
Option B: Capture Cookies via Wget/Curl (Programmatic Login)
If you haven't logged in yet, use wget to send a login POST request and save the cookies:
wget --save-cookies cookies.txt --keep-session-cookies --post-data "username=your_username&password=your_password" https://your-private-site.com/login-page
The --keep-session-cookies flag ensures temporary session cookies (without an expiration date) are saved too—critical for most private sites.
Once you have your cookies.txt file and your 2000 links saved in links.txt (one link per line), run this command:
wget --load-cookies cookies.txt --keep-session-cookies -i links.txt
Key flag breakdown:
--load-cookies cookies.txt: Tells wget to use your saved cookies for every request.--keep-session-cookies: Retains session cookies between requests to avoid unexpected logouts mid-batch.-i links.txt: Reads all URLs from your text file instead of specifying them one by one.
Bonus: Add Rate Limiting to Avoid Blocking
Since you're sending 2000 requests, add these flags to be polite and prevent IP bans:
wget --load-cookies cookies.txt --keep-session-cookies -i links.txt --wait=2 --limit-rate=100k --retry-connrefused --continue
--wait=2: Waits 2 seconds between each request.--limit-rate=100k: Caps download speed at 100KB/s.--retry-connrefused: Retries if the site temporarily refuses your connection.--continue: Resumes partially downloaded files if a request fails mid-download.
If you prefer curl (or hit issues with wget), here's how to replicate the workflow:
Ensure your cookie file is in Netscape format, then run:
xargs -a links.txt -P 5 curl -b cookies.txt -O --limit-rate 100k --retry 3
-a links.txt: Reads URLs from your list file.-P 5: Runs up to 5 concurrent requests (adjust based on the site's strictness).-b cookies.txt: Loads cookies from your saved file.-O: Saves each file with its original filename from the URL.--limit-rate 100k/--retry 3: Limits speed and retries failed requests.
- Cookie Expiration: Session cookies often expire after hours or days. If your batch takes a long time, re-export cookies or re-run the login command halfway through.
- Cookie Scope: Double-check that your cookie file includes all relevant entries (domain and path must match your target links).
- Security: Keep your
cookies.txtfile private (usechmod 600on Linux/macOS) since it contains sensitive login data.
内容的提问来源于stack exchange,提问作者Josh

