Python网页爬取:爬取CoinMarketCap时如何跳过无效URL?
How to Skip Invalid URLs When Scraping CoinMarketCap
Hey there! It sounds like you're working on pulling historical crypto data from CoinMarketCap but hitting a snag with broken URLs. No worries—we can fix this by adding error handling to your code to catch those invalid links and skip them gracefully.
Here's how to adjust your code:
The core problem is that urllib.request.urlopen() will throw an error (like HTTPError for 404 pages) when it runs into an invalid URL. We can wrap the URL fetching and parsing logic in a try-except block to catch these errors and move on to the next currency without breaking the whole loop.
import urllib.request from bs4 import BeautifulSoup from urllib.error import HTTPError, URLError # Assuming `name` and `name_size` are defined in your code for j in range(0, name_size): coin_name = str(name[j]) url = f"https://coinmarketcap.com/currencies/{coin_name}/historical-data/?start=20130429&end=20191016" try: # Attempt to open the URL and parse the page page = urllib.request.urlopen(url) soup = BeautifulSoup(page, 'html.parser') priceDiv = soup.find('div', attrs={'class':'table-responsive'}) rows = priceDiv.find_all('tr') # Add your code here to extract Open, High, Low, Close, Marketcap, Volume from rows print(f"Successfully scraped data for {coin_name}") except HTTPError as e: # Catch cases where the URL returns an error (like 404 Not Found) print(f"Skipping {coin_name}: HTTP Error {e.code} for URL {url}") continue # Move to the next currency except URLError as e: # Handle connection issues (e.g., no internet, invalid domain) print(f"Skipping {coin_name}: Failed to connect to {url} - {e.reason}") continue except AttributeError: # Cover cases where the page loads but the data table doesn't exist (priceDiv is None) print(f"Skipping {coin_name}: Data table not found on page") continue
Key improvements:
- HTTPError handling: Catches invalid URLs that return 404/500 errors and skips them instead of crashing the script.
- URLError handling: Deals with connection problems like lost internet or malformed domain names.
- AttributeError handling: Prevents crashes if the page loads but the expected data table is missing (so
priceDivwould beNone). - Debug print statements: Lets you track which coins failed and why, making it easier to troubleshoot later.
Extra tips:
- Add a small delay between requests (using
time.sleep(1)or similar) to avoid triggering CoinMarketCap's anti-scraping measures. - Consider storing the names of coins with invalid URLs in a separate list—this can help you spot typos in your
namearray or coins that didn't exist during your target date range.
内容的提问来源于stack exchange,提问作者studyrivulet
相关产品推荐
相关产品推荐

