如何用Python的urllib或requests爬取goldprice.org的最新黄金兑美元汇率
How to Scrape the Latest Gold-to-USD Exchange Rate from goldprice.org Using Python
Great question! Let's walk through how to get that latest gold price using either urllib (building on your sample code) or the more intuitive requests library. I'll also fix up your example to actually extract the clean price text instead of just returning a list of elements.
Using urllib (Your Original Approach, Refined)
Your initial code correctly targets the price element, but it returns a list of elements instead of the actual price string. Here's the improved version:
from urllib.request import urlopen from bs4 import BeautifulSoup # Fetch the webpage html = urlopen("https://goldprice.org/") bs = BeautifulSoup(html, "html.parser") # Find the unique price element (use find() instead of findAll() for efficiency) price_element = bs.find(id="gpxtickerLeft_price") # Extract and clean the price text if the element exists if price_element: gold_price_usd = price_element.get_text(strip=True) print(f"Latest Gold Price (USD): {gold_price_usd}") else: print("Couldn't locate the gold price element—double-check if the page's HTML structure has changed!")
Key Improvements:
- Swapped
findAll()forfind(): SincegpxtickerLeft_priceis a unique ID,find()stops at the first match (more efficient than returning a list). - Added
get_text(strip=True): This removes extra whitespace and pulls out the clean price value from the element.
Using requests (Recommended for Simplicity)
The requests library simplifies HTTP requests and handles common edge cases better than raw urllib. Here's how to implement the same scrape with it:
First, install requests if you haven't already:
pip install requests
Then the code:
import requests from bs4 import BeautifulSoup # Fetch the page with requests (raises an error if the request fails) response = requests.get("https://goldprice.org/") response.raise_for_status() # Parse the HTML content bs = BeautifulSoup(response.text, "html.parser") # Extract the price price_element = bs.find(id="gpxtickerLeft_price") if price_element: gold_price_usd = price_element.get_text(strip=True) print(f"Latest Gold Price (USD): {gold_price_usd}") else: print("Price element not found—verify the page's HTML structure hasn't been updated.")
Important Considerations:
- Page Structure Changes: Websites often update their HTML. If this code stops working, right-click the price on goldprice.org and select "Inspect" to confirm the element's ID is still
gpxtickerLeft_price. - Respect Site Policies: Check the site's
robots.txtfile to ensure scraping is allowed, and avoid making too many requests in a short time to prevent getting blocked. - JavaScript-Rendered Content: As of now, the gold price is in static HTML, but if the site ever switches to loading prices via JavaScript, you'll need tools like
seleniumorplaywrightto render the page fully.
内容的提问来源于stack exchange,提问作者dauhuong68
相关产品推荐
相关产品推荐

