如何用Python抓取网站特定文本:提取指定在线玩家数并存入变量
Got it, let's break down how to build this Python scraper step by step. The goal is to pull that specific "X Players Online" text and store it in a variable you can use later—here's how to do it:
First, you'll need two essential tools for this job:
requests: Fetches the raw HTML content from the websiteBeautifulSoup: Parses the messy HTML into something we can easily search through
Install them via pip if you haven't already:
pip install requests beautifulsoup4
Here's a complete, commented example that targets the Growtopia website (we’ll adjust based on how their page is structured):
import requests from bs4 import BeautifulSoup # Define the target URL url = "https://growtopiagame.com/" try: # Send a request to the website (simulate a browser to avoid being blocked) headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36" } response = requests.get(url, headers=headers) response.raise_for_status() # Raise an error if the request fails (e.g., 404, 500) # Parse the HTML content soup = BeautifulSoup(response.text, "html.parser") # Find the element containing the "Players Online" text # IMPORTANT: Inspect the website's HTML to find the right selector! # For Growtopia, let's say the text is inside a div with id "player-count" (check page source to confirm) online_element = soup.find("div", id="player-count") if online_element: # Extract the text and clean up extra whitespace online_players_text = online_element.get_text(strip=True) # Now you have your variable ready for later use! print(f"Fetched: {online_players_text}") # Optional: Extract just the number if you need it player_count = online_players_text.split()[0] print(f"Just the player number: {player_count}") else: print("Couldn't find the 'Players Online' element—double-check the page structure!") except requests.exceptions.RequestException as e: print(f"Error fetching the page: {e}")
- Inspect the Website's HTML: Right-click the "X Players Online" text on the page and select "Inspect" to see what HTML tag/class/id wraps it. Adjust the
soup.find()parameters to match (e.g., if it’s a span with class "online-users", usesoup.find("span", class_="online-users")). - Avoid Being Blocked: Adding a
User-Agentheader makes your request look like it’s coming from a real browser, not a bot—many sites block requests without this. - Handle Errors: The
try-exceptblock catches issues like broken links, network errors, or missing elements so your program doesn’t crash unexpectedly.
Once you run this, the online_players_text variable will hold the full "4231 Players Online" string, and you can use it anywhere else in your code—whether that’s saving it to a file, displaying it in an app, or doing further calculations.
内容的提问来源于stack exchange,提问作者Well Dirth

