You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup爬取遇阻:无法提取巴甲联赛球队列表

Troubleshooting Your Brasileirão Table Scraper

Hey Lucas, let's break down why you're not getting any team data from your web scrape, and fix it up. From your code and screenshots, here are the key issues and actionable fixes:

Common Issues Blocking Your Scrape

  • Bot Detection: By default, the requests library sends a User-Agent that tells Google you're a script, not a real browser. Google often serves a stripped-down or blocked page to bots, so your target elements won't exist.
  • Dynamic Class Names: Google's HTML classes (like ellipsisize hsKSJe) are dynamically generated and change frequently. The class you're targeting might already be outdated by the time you run your code.
  • JS-Rendered Content: Some parts of Google's search results load after the initial page load via JavaScript. requests only fetches static HTML, so it can't access content that loads later.

Fix 1: Add Browser-like Request Headers

First, let's bypass bot detection by adding a valid User-Agent to mimic a real browser. This should get you the full search results page:

from bs4 import BeautifulSoup
import requests

# Mimic a Chrome browser request
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36"
}

# Cleaned up URL (removed the unnecessary #sie fragment)
url = "https://www.google.com/search?q=Tabela+do+Campeonato+Brasileiro+de+Futebol&oq=Tabela+do+Campeonato+Brasileiro+de+Futebol&aqs=chrome..69i57.241j0j1&sourceid=chrome&ie=UTF-8"

try:
    r = requests.get(url, headers=headers)
    r.raise_for_status()  # Raise an error if the request fails (e.g., 403 Forbidden)
    soup = BeautifulSoup(r.text, "lxml")

    # Target the table container first (more stable than individual classes)
    table_container = soup.find("div", class_="sports-featured-table")
    if table_container:
        # Use a partial class match to avoid breaking if Google changes class names
        team_spans = table_container.find_all("span", class_=lambda x: x and "ellipsisize" in x)
        for team in team_spans:
            print(team.get_text(strip=True))
    else:
        print("Couldn't find the Brasileirão table container – Google's page structure may have changed.")
except Exception as e:
    print(f"Error during scrape: {str(e)}")

Fix 2: Use Selenium for Dynamic Content

If the table still doesn't load (because it's rendered with JavaScript), use Selenium to simulate a real browser that executes JS:

First, install Selenium:

pip install selenium

Then run this code:

from bs4 import BeautifulSoup
from selenium import webdriver
from selenium.webdriver.chrome.options import Options
import time

# Configure Chrome to run in headless mode (no visible window)
chrome_options = Options()
chrome_options.add_argument("--headless=new")
chrome_options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36")

driver = webdriver.Chrome(options=chrome_options)
url = "https://www.google.com/search?q=Tabela+do+Campeonato+Brasileiro+de+Futebol&oq=Tabela+do+Campeonato+Brasileiro+de+Futebol&aqs=chrome..69i57.241j0j1&sourceid=chrome&ie=UTF-8"

try:
    driver.get(url)
    time.sleep(2)  # Give JavaScript time to load the table

    soup = BeautifulSoup(driver.page_source, "lxml")
    team_spans = soup.find_all("span", class_="ellipsisize hsKSJe")
    
    if team_spans:
        for team in team_spans:
            print(team.get_text(strip=True))
    else:
        print("Team elements not found – check if the class name has changed.")
finally:
    driver.quit()  # Always close the browser to avoid memory leaks

Pro Tips for Reliable Scraping

  • Avoid Google for Sports Data: Google's page structure changes constantly. Instead, scrape dedicated sports sites (like Globo Esporte's Brasileirão page) – their HTML is more stable, and they're less likely to block bots.
  • Use Partial Class Matches: Instead of targeting full class names, use lambda x: x and "ellipsisize" in x to match any span with "ellipsisize" in its class.
  • Rate Limiting: Add delays between requests to avoid getting your IP blocked by Google.

内容的提问来源于stack exchange,提问作者Lucas Ribeiro

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 12:57:50