如何程序化登录HTTP网站并像FTP一样上传/下载文件?
Absolutely—you can totally automate logging into that HTTP site and handle file uploads/downloads even without an official API. I’ve helped folks solve similar FTP fallback issues before, so here’s a practical breakdown of how to pull this off:
1. First, Map Out the Site’s Manual Workflow
Before writing any code, you need to reverse-engineer exactly what happens when you log in and transfer files manually:
- Fire up your browser’s dev tools (F12 → Network tab) and record every request while you:
- Load the login page
- Submit your credentials
- Upload a file
- Download a file
- Jot down key details:
- The POST URL for login, plus the exact form field names (e.g.,
username,password,csrf_token) - How the site maintains your session (usually via cookies or an auth header)
- The upload endpoint, and the name attribute of the file input (e.g.,
<input type="file" name="user_file">) - The direct download URL for files, or how the site triggers downloads
- The POST URL for login, plus the exact form field names (e.g.,
2. Pick Your Tooling (By Language)
Choose a library that fits your preferred stack—here are the most common options:
Python (Most Flexible)
Use requests for session management and requests-toolbelt for multipart file uploads. Here’s a stripped-down example:
import requests from requests_toolbelt.multipart.encoder import MultipartEncoder # Initialize a session to automatically persist cookies session = requests.Session() # Step 1: Grab any required CSRF token (skip if the site doesn't use one) login_page = session.get("https://your-target-site.com/login") # Use BeautifulSoup to parse the token from the page HTML if needed # from bs4 import BeautifulSoup # soup = BeautifulSoup(login_page.text, "html.parser") # csrf_token = soup.find("input", {"name": "csrf_token"})["value"] # Step 2: Log in login_payload = { "username": "your-account", "password": "your-password", # "csrf_token": csrf_token # Uncomment if you found a token } login_response = session.post("https://your-target-site.com/login", data=login_payload) # Verify login success (adjust this check to match the site's behavior) if "Welcome" in login_response.text and login_response.status_code == 200: print("Login successful!") # Step 3: Upload a file upload_data = MultipartEncoder( fields={ "user_file": ("report.pdf", open("report.pdf", "rb"), "application/pdf"), "target_folder": "456" # Add any extra form fields the site requires } ) upload_response = session.post( "https://your-target-site.com/upload", data=upload_data, headers={"Content-Type": upload_data.content_type} ) print(f"Upload result: {upload_response.status_code}") # Step 4: Download a file download_response = session.get("https://your-target-site.com/download/789") with open("downloaded_report.pdf", "wb") as f: f.write(download_response.content)
Shell Scripts
If you prefer no-code/low-code, use curl with cookie persistence:
# Save cookies to a file after login curl -c cookies.txt -d "username=your-account&password=your-password" https://your-target-site.com/login # Upload a file using saved cookies curl -b cookies.txt -F "user_file=@report.pdf" -F "target_folder=456" https://your-target-site.com/upload # Download a file curl -b cookies.txt -o downloaded_report.pdf https://your-target-site.com/download/789
3. Watch Out for Common Pitfalls
- CSRF Tokens: Most modern sites use these to block automated requests. Always grab the token from the login page before submitting credentials.
- Session Expiry: Your logged-in session might time out—add logic to check if the session is still valid, and re-login automatically if needed.
- Anti-Bot Measures: If the site uses captchas or rate limits, you might need to add delays, use a headless browser, or even reach out to the site owner for exceptions (since you’re just replacing a broken FTP connection).
- Page Changes: If the site updates its HTML structure or endpoints, your script will break. Schedule occasional checks to verify everything still works.
4. Alternative: Headless Browser Automation
If reverse-engineering requests feels too tedious, use tools like Playwright or Selenium to simulate real browser actions (clicks, typing, file selection). This is more robust against site changes but uses more resources:
from playwright.sync_api import sync_playwright with sync_playwright() as p: browser = p.chromium.launch(headless=True) # Set to False to see the browser page = browser.new_page() # Log in page.goto("https://your-target-site.com/login") page.fill("#username-input", "your-account") page.fill("#password-input", "your-password") page.click("button[type='submit']") page.wait_for_url("https://your-target-site.com/dashboard") # Upload file page.click("#upload-button") page.set_input_files("#file-upload-input", "report.pdf") page.click("#confirm-upload") # Download file with page.expect_download() as download_info: page.click("#download-report-button") download = download_info.value download.save_as("downloaded_report.pdf") browser.close()
内容的提问来源于stack exchange,提问作者Brian Evans

