使用BeautifulSoup爬取遇阻:无法提取巴甲联赛球队列表
Troubleshooting Your Brasileirão Table Scraper
Hey Lucas, let's break down why you're not getting any team data from your web scrape, and fix it up. From your code and screenshots, here are the key issues and actionable fixes:
Common Issues Blocking Your Scrape
- Bot Detection: By default, the
requestslibrary sends aUser-Agentthat tells Google you're a script, not a real browser. Google often serves a stripped-down or blocked page to bots, so your target elements won't exist. - Dynamic Class Names: Google's HTML classes (like
ellipsisize hsKSJe) are dynamically generated and change frequently. The class you're targeting might already be outdated by the time you run your code. - JS-Rendered Content: Some parts of Google's search results load after the initial page load via JavaScript.
requestsonly fetches static HTML, so it can't access content that loads later.
Fix 1: Add Browser-like Request Headers
First, let's bypass bot detection by adding a valid User-Agent to mimic a real browser. This should get you the full search results page:
from bs4 import BeautifulSoup import requests # Mimic a Chrome browser request headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36" } # Cleaned up URL (removed the unnecessary #sie fragment) url = "https://www.google.com/search?q=Tabela+do+Campeonato+Brasileiro+de+Futebol&oq=Tabela+do+Campeonato+Brasileiro+de+Futebol&aqs=chrome..69i57.241j0j1&sourceid=chrome&ie=UTF-8" try: r = requests.get(url, headers=headers) r.raise_for_status() # Raise an error if the request fails (e.g., 403 Forbidden) soup = BeautifulSoup(r.text, "lxml") # Target the table container first (more stable than individual classes) table_container = soup.find("div", class_="sports-featured-table") if table_container: # Use a partial class match to avoid breaking if Google changes class names team_spans = table_container.find_all("span", class_=lambda x: x and "ellipsisize" in x) for team in team_spans: print(team.get_text(strip=True)) else: print("Couldn't find the Brasileirão table container – Google's page structure may have changed.") except Exception as e: print(f"Error during scrape: {str(e)}")
Fix 2: Use Selenium for Dynamic Content
If the table still doesn't load (because it's rendered with JavaScript), use Selenium to simulate a real browser that executes JS:
First, install Selenium:
pip install selenium
Then run this code:
from bs4 import BeautifulSoup from selenium import webdriver from selenium.webdriver.chrome.options import Options import time # Configure Chrome to run in headless mode (no visible window) chrome_options = Options() chrome_options.add_argument("--headless=new") chrome_options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36") driver = webdriver.Chrome(options=chrome_options) url = "https://www.google.com/search?q=Tabela+do+Campeonato+Brasileiro+de+Futebol&oq=Tabela+do+Campeonato+Brasileiro+de+Futebol&aqs=chrome..69i57.241j0j1&sourceid=chrome&ie=UTF-8" try: driver.get(url) time.sleep(2) # Give JavaScript time to load the table soup = BeautifulSoup(driver.page_source, "lxml") team_spans = soup.find_all("span", class_="ellipsisize hsKSJe") if team_spans: for team in team_spans: print(team.get_text(strip=True)) else: print("Team elements not found – check if the class name has changed.") finally: driver.quit() # Always close the browser to avoid memory leaks
Pro Tips for Reliable Scraping
- Avoid Google for Sports Data: Google's page structure changes constantly. Instead, scrape dedicated sports sites (like Globo Esporte's Brasileirão page) – their HTML is more stable, and they're less likely to block bots.
- Use Partial Class Matches: Instead of targeting full class names, use
lambda x: x and "ellipsisize" in xto match any span with "ellipsisize" in its class. - Rate Limiting: Add delays between requests to avoid getting your IP blocked by Google.
内容的提问来源于stack exchange,提问作者Lucas Ribeiro
相关产品推荐
相关产品推荐

