如何用Python(非Selenium)抓取http://shinytoylabs.com/jargon/生成文本?
Scrape Jargon Text from shinytoylabs.com/jargon/ with Requests (No Selenium)
Got it, I’ve run into this exact kind of dynamic content issue before—don’t worry, we don’t need Selenium here. The button click on that page is just triggering an asynchronous API call under the hood, so we can skip rendering the page entirely and hit that API directly with requests.
Step 1: Find the Hidden API Endpoint
First, let’s figure out what’s happening when you click the button:
- Open your browser’s DevTools (F12 or Ctrl+Shift+I), go to the Network tab, and filter for "Fetch/XHR".
- Click the "Generate Jargon" button on the page. You’ll see a new POST request pop up to
http://shinytoylabs.com/jargon/generate. - Check the response: it’s a JSON object that contains the generated jargon text right there (look for the
jargonkey).
Step 2: Code to Fetch the Text with Requests
Now we can replicate that API call in Python. Here’s a complete working example:
import requests def get_jargon(): # The API endpoint we found api_url = "http://shinytoylabs.com/jargon/generate" # Add a basic User-Agent to avoid being blocked (many sites reject default requests) headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36" } try: # Send the POST request (just like the browser does) response = requests.post(api_url, headers=headers) response.raise_for_status() # Raise an error if the request fails # Parse the JSON response jargon_data = response.json() # Extract the generated jargon text return jargon_data.get("jargon", "Couldn't retrieve jargon text.") except requests.exceptions.RequestException as e: return f"Request failed: {str(e)}" # Test it out if __name__ == "__main__": print(get_jargon())
Why This Works
Instead of simulating a button click and waiting for the page to render, we’re directly calling the backend API that generates the text. This is faster and more reliable than using Selenium for cases like this.
Notes
- If you run into issues with being blocked, you can add more headers from the browser’s request (like
RefererorAccept), but the User-Agent should be enough for this site. - The API doesn’t require any POST parameters—just a simple POST request to the endpoint is all it takes.
内容的提问来源于stack exchange,提问作者miamiamia
相关产品推荐
相关产品推荐

