如何用BeautifulSoup从panel-body类div提取Current Calculated Hashrate数值
Hey there! Let's work through getting that "Current Calculated Hashrate" value extracted properly. Since your initial attempt fell flat, let's break down common scenarios and fixes based on how the page is structured:
1. If the page uses static HTML (no JavaScript rendering)
Chances are you might be targeting the wrong nested element, or grabbing too many panel-body divs by mistake. Let's narrow down the panel first using its title, then pull the value:
from bs4 import BeautifulSoup import requests # Replace with your target URL target_url = "your_link_here" # Add a user-agent to avoid being blocked as a bot headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } response = requests.get(target_url, headers=headers) soup = BeautifulSoup(response.text, "html.parser") # First, find the entire panel that contains the "Current Calculated Hashrate" title target_panel = soup.find( "div", class_="panel", # Look for the title text inside the panel (adjust if the title is in a <h3> or <div> child) string=lambda text: text and "Current Calculated Hashrate" in text.strip() ) if target_panel: # Grab the panel-body inside this specific panel panel_body = target_panel.find("div", class_="panel-body") if panel_body: # Extract raw text, then clean it to get just the numeric value hashrate_text = panel_body.get_text(strip=True) # Filter out non-digit characters (adjust if you need to keep decimals/units) hashrate_value = "".join(filter(str.isdigit, hashrate_text)) print(f"Current Calculated Hashrate: {hashrate_value}") else: print("Couldn't locate the panel-body inside the target panel") else: print("Couldn't find the panel with 'Current Calculated Hashrate' title")
2. If the page uses JavaScript to load the hashrate
If the value only appears after the page runs JS (you won't see it in the raw HTML source), you'll need a tool that can render the page like a browser. Selenium works great for this:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC target_url = "your_link_here" # Initialize Chrome driver (make sure you have the ChromeDriver installed) driver = webdriver.Chrome() try: driver.get(target_url) # Wait up to 10 seconds for the target panel to load target_panel = WebDriverWait(driver, 10).until( EC.presence_of_element_located( (By.XPATH, "//div[@class='panel' and contains(., 'Current Calculated Hashrate')]") ) ) # Grab the panel-body and extract the value panel_body = target_panel.find_element(By.CLASS_NAME, "panel-body") hashrate_text = panel_body.text.strip() hashrate_value = "".join(filter(str.isdigit, hashrate_text)) print(f"Current Calculated Hashrate: {hashrate_value}") finally: # Always close the browser when done driver.quit()
Quick Troubleshooting Tips
- Double-check the HTML structure: Right-click the hashrate element in your browser, select "Inspect" to confirm exactly where the value lives. It might be nested inside a
<span>or<p>tag within thepanel-body. - Watch for anti-scraping measures: Some sites block requests without a proper user-agent, or require a logged-in session. If you're getting empty responses, try adding cookies from your browser to the request.
内容的提问来源于stack exchange,提问作者T.T
相关产品推荐
相关产品推荐

