使用requests与Bs4抓取Yourupload自动生成直链并保存至文件
Hey there! Let's figure out how to grab that auto-generated direct download link from Yourupload and save it to Link.txt. Your current code just prints the raw page soup, so we need to dig into how the site serves the download link after the 5-second wait.
First, a quick heads-up: Yourupload uses JavaScript to handle the countdown and generate the direct link. Since requests doesn't execute JavaScript, we'll need to either parse the static page for hidden links/parameters or mimic the API calls the site makes once the countdown ends.
Here's a step-by-step solution with working code:
Step 1: Mimic a browser request
Most sites block requests without a proper User-Agent header, so we'll add that to avoid being flagged as a bot.
Step 2: Extract the download link
We'll try a few approaches (since the site's structure might shift over time):
- Hunt for visible/hidden download links in the page's HTML
- Parse embedded JavaScript code for the redirect URL
- Mimic the API call the site uses to fetch the final link
Full Code Implementation
import requests from bs4 import BeautifulSoup import re # Target download page URL target_url = "https://www.yourupload.com/download?file=2573285" # Headers to mimic a modern Chrome browser request headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } # Fetch the initial page response = requests.get(target_url, headers=headers) soup = BeautifulSoup(response.text, "lxml") direct_link = None # Approach 1: Check anchor tags for direct download links for anchor in soup.find_all("a", href=True): link_href = anchor["href"] if "dl.yourupload" in link_href or "download" in link_href.lower(): direct_link = link_href break # Approach 2: Extract link from JavaScript redirect code if not direct_link: for script in soup.find_all("script"): script_content = script.string if script_content and ("window.location" in script_content or "downloadUrl" in script_content): # Use regex to pull the URL from the JS code url_match = re.search(r'["\'](https?://[^"\']+download[^"\']+)["\']', script_content) if url_match: direct_link = url_match.group(1) break # Approach 3: Fetch link via the site's internal API if not direct_link: # Look for hidden form tokens common on download sites token_input = soup.find("input", {"name": "token"}) or soup.find("input", {"name": "file_token"}) if token_input: token_value = token_input["value"] # Construct the API request (adjust endpoint if needed - check your browser's network tab) api_endpoint = "https://www.yourupload.com/api/get_download" payload = {"file": "2573285", "token": token_value} api_response = requests.post(api_endpoint, headers=headers, data=payload) if api_response.ok: # Assume the API returns JSON with the download link direct_link = api_response.json().get("url") # Save the link to Link.txt if we found it if direct_link: # Fix relative paths to make the link absolute if not direct_link.startswith("http"): direct_link = f"https://www.yourupload.com{direct_link}" with open("Link.txt", "w", encoding="utf-8") as file: file.write(direct_link) print(f"Success! Direct download link saved to Link.txt:\n{direct_link}") else: print("Couldn't find the direct download link. The site's structure may have changed - check the page source or browser network tab for updates.")
Troubleshooting Tips
- If none of the above work, open the download page in Chrome/Firefox, launch the Network tab in DevTools, wait for the 5-second countdown, and look for the request that loads the direct download link. You can then replicate that exact request in your code.
- Always keep your
User-Agentupdated to match a modern browser to avoid being blocked.
内容的提问来源于stack exchange,提问作者user11676469

