如何用Cheerio从同一<tr>标签中提取多个体育赔率数据
Hey there! I see you're building a web scraper to collect MLB odds data for your personal small database—let's walk through how to extract the exact info you need from that page's Bets section.
First, Let's Break Down the HTML Structure
From the snippet you shared, each <tr> row contains all the odds-related data we care about:
- The
<td>with classtext-right border-leftholds the percentage values (like51%and49%) <span>elements with classhighlight-greencontain the odds (e.g.,-130) and point spreads (e.g.,9.5)- Other
<td>s have subscription-locked content, but we can ignore those since we're targeting the visible data
Tools to Use
For a small personal project, Python's requests + BeautifulSoup is perfect—they're lightweight, easy to set up, and handle static HTML well. If the page loads content dynamically with JavaScript later, you can switch to Selenium, but let's start with the basics.
Step-by-Step Implementation
1. Install Required Libraries
First, install the dependencies via pip:
pip install requests beautifulsoup4
2. Fetch the Page Content
Most sites block bare requests calls, so we'll add a user-agent header to mimic a browser:
import requests from bs4 import BeautifulSoup # Target URL url = "https://www.actionnetwork.com/mlb/live-odds" # Mimic a browser request headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } # Get the page and parse it with BeautifulSoup response = requests.get(url, headers=headers) soup = BeautifulSoup(response.text, "html.parser")
3. Extract Data from Each Row
Loop through every <tr> in the target <tbody> and pull the data you need:
# Get all rows from the table body rows = soup.select("tbody tr") for row in rows: # Extract percentage data from the border-left TD percentage_cell = row.select_one("td.text-right.border-left") if percentage_cell: percentages = [span.text.strip() for span in percentage_cell.find_all("span", class_="d-block")] print(f"Percentage Split: {percentages}") # Extract odds and spread values from highlight-green spans odds_spans = row.find_all("span", class_="highlight-green") # Odds and spreads are paired (odds first, then spread) odds_spread_pairs = [] for i in range(0, len(odds_spans), 2): if i + 1 < len(odds_spans): odds = odds_spans[i].text.strip() spread = odds_spans[i+1].text.strip() odds_spread_pairs.append((odds, spread)) print(f"Odds & Spreads: {odds_spread_pairs}") # Optional: Extract the "No Picks" status if needed no_picks_cell = row.select_one("td.text-right.border-left + td") if no_picks_cell: picks_status = no_picks_cell.text.strip() print(f"Picks Status: {picks_status}") # Add a separator between rows for readability print("---")
Key Notes to Keep in Mind
- Anti-Scraping Measures: If you get a 403 Forbidden error or empty content, the site might be blocking your request. Try adding more headers (like
Accept-Language) or usingSeleniumto simulate a real browser session. - HTML Structure Changes: Websites often update their markup—keep an eye on your selectors and adjust them if the site changes.
- Compliance: Make sure your scraping follows the site's
robots.txtand terms of service to avoid any legal issues.
内容的提问来源于stack exchange,提问作者G. Boyce

