You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python网页爬取:爬取CoinMarketCap时如何跳过无效URL?

How to Skip Invalid URLs When Scraping CoinMarketCap

Hey there! It sounds like you're working on pulling historical crypto data from CoinMarketCap but hitting a snag with broken URLs. No worries—we can fix this by adding error handling to your code to catch those invalid links and skip them gracefully.

Here's how to adjust your code:

The core problem is that urllib.request.urlopen() will throw an error (like HTTPError for 404 pages) when it runs into an invalid URL. We can wrap the URL fetching and parsing logic in a try-except block to catch these errors and move on to the next currency without breaking the whole loop.

import urllib.request
from bs4 import BeautifulSoup
from urllib.error import HTTPError, URLError

# Assuming `name` and `name_size` are defined in your code
for j in range(0, name_size):
    coin_name = str(name[j])
    url = f"https://coinmarketcap.com/currencies/{coin_name}/historical-data/?start=20130429&end=20191016"
    
    try:
        # Attempt to open the URL and parse the page
        page = urllib.request.urlopen(url)
        soup = BeautifulSoup(page, 'html.parser')
        priceDiv = soup.find('div', attrs={'class':'table-responsive'})
        rows = priceDiv.find_all('tr')
        
        # Add your code here to extract Open, High, Low, Close, Marketcap, Volume from rows
        print(f"Successfully scraped data for {coin_name}")
        
    except HTTPError as e:
        # Catch cases where the URL returns an error (like 404 Not Found)
        print(f"Skipping {coin_name}: HTTP Error {e.code} for URL {url}")
        continue  # Move to the next currency
    except URLError as e:
        # Handle connection issues (e.g., no internet, invalid domain)
        print(f"Skipping {coin_name}: Failed to connect to {url} - {e.reason}")
        continue
    except AttributeError:
        # Cover cases where the page loads but the data table doesn't exist (priceDiv is None)
        print(f"Skipping {coin_name}: Data table not found on page")
        continue

Key improvements:

  • HTTPError handling: Catches invalid URLs that return 404/500 errors and skips them instead of crashing the script.
  • URLError handling: Deals with connection problems like lost internet or malformed domain names.
  • AttributeError handling: Prevents crashes if the page loads but the expected data table is missing (so priceDiv would be None).
  • Debug print statements: Lets you track which coins failed and why, making it easier to troubleshoot later.

Extra tips:

  • Add a small delay between requests (using time.sleep(1) or similar) to avoid triggering CoinMarketCap's anti-scraping measures.
  • Consider storing the names of coins with invalid URLs in a separate list—this can help you spot typos in your name array or coins that didn't exist during your target date range.

内容的提问来源于stack exchange,提问作者studyrivulet

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:19:15