You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从含标签内ID的HTML字符串提取数据?BeautifulSoup求助

Solution to Extract Target Race Data with BeautifulSoup

Got it, let's work through your problem step by step. You're hitting two key issues: escaped HTML entities that mess up BeautifulSoup's parsing, and extracting embedded data (like IDs from href attributes) alongside visible text. Here's how to fix both:

1. Fix HTML Entity Escaping First

Your raw HTML has escaped characters like <, >, and & that prevent BeautifulSoup from recognizing proper tags. We'll use Python's built-in html.unescape() to convert these back to valid HTML syntax first.

2. Extract Data with Targeted Parsing

Once the HTML is properly unescaped, we can iterate through each race row, pull out both attribute data (like r_ID and w_ID from href links) and visible text values.

Full Working Code Example

from bs4 import BeautifulSoup
import html

# Your raw HTML string (from your input)
raw_html = """
<th><a href="d?racename=&amp;country=1000&amp;startmonth=1&amp;endmonth=10&amp;startdate=2018&amp;enddate=2019&amp;maxdist=unlimitied&amp;class=any&amp;x=1&amp;order=winner&amp;z=Px_8iD">Winner</a> </th> <th background="b8.gif" width="30" title="Winning time - click on this header to sort results by this column"> <a href="d?racename=&amp;country=1000&amp;startmonth=1&amp;endmonth=10&amp;startdate=2018&amp;enddate=2019&amp;maxdist=unlimitied&amp;class=any&amp;x=1&amp;order=wintime&amp;z=Px_8iD">Wintime</a> </th> <th background="b8.gif" title="races with icon have video available for download">Film</th> </tr>
<tr> <td><a href="d?r=4552510&amp;z=Px_8iD">OAKS AT LOGAN PARK (1-2 WINS)</a></td> <td>Warragul</td> <td>18;OCT;2019</td> <td>7</td> <td>GR;Tier</td> <td>460;503</td> <td><a href="d?i=2390975">Madalia Ken</a></td> <td>26.00</td> <td></td> </tr>
<tr bgcolor="#cccccc"> <td><a href="d?r=4552511&amp;z=Px_8iD">AUSTRALIAN QUALITY PET FOODS</a></td> <td>Warragul</td> <td>18;OCT;2019</td> <td>8</td> <td>GR;Grad</td> <td>460;503</td> <td><a href="d?i=2304665">Midnight Storm</a></td> <td>26.24</td> <td></td> </tr>
<tr> <td><a href="d?r=4552512&amp;z=Px_8iD">EAST IVANHOE GROCERS</a></td> <td>Warragul</td> <td>18;OCT;2019</td> <td>9</td> <td>GR;Grad</td> <td>400;437</td> <td><a href="d?i=2362422">Early Promise</a></td> <td>23.15</td> <td></td> </tr>
"""

# Step 1: Unescape HTML entities to fix broken tags
unescaped_html = html.unescape(raw_html)

# Step 2: Parse with BeautifulSoup
soup = BeautifulSoup(unescaped_html, 'html.parser')

# Step 3: Extract race rows (skip the first row which is the header)
race_rows = soup.find_all('tr')[1:]

# Step 4: Iterate through each row and extract target fields
extracted_data = []
for row in race_rows:
    tds = row.find_all('td')
    
    # Extract race ID and name from first td's link
    race_link = tds[0].find('a')
    r_id = race_link['href'].split('r=')[1].split('&')[0]
    r_name = race_link.get_text(strip=True)
    
    # Extract winner ID and name from 7th td's link
    winner_link = tds[6].find('a')
    w_id = winner_link['href'].split('i=')[1]
    w_name = winner_link.get_text(strip=True)
    
    # Compile all fields into a dictionary
    row_data = {
        "r_ID": r_id,
        "r_name": r_name,
        "s_name": tds[1].get_text(strip=True),
        "date": tds[2].get_text(strip=True),
        "h": tds[3].get_text(strip=True),
        "g": tds[4].get_text(strip=True),
        "d": tds[5].get_text(strip=True),
        "w_ID": w_id,
        "w_name": w_name,
        "wt": tds[7].get_text(strip=True),
        "f": tds[8].get_text(strip=True) or ""
    }
    extracted_data.append(row_data)

# Print or use the extracted data
for idx, data in enumerate(extracted_data, 1):
    print(f"Row {idx}:")
    for key, value in data.items():
        print(f"  {key}: {value}")

Key Fixes Explained

  • HTML Unescaping: html.unescape() converts entities like &amp;gt; back to >, so BeautifulSoup can correctly identify <a> and <td> tags instead of treating them as plain text.
  • Attribute Data Extraction: We pull IDs from the href attribute by splitting the URL string (e.g., split('r=')[1].split('&')[0] isolates the race ID from the link parameters).
  • Robust Text Extraction: get_text(strip=True) removes extra whitespace from visible text values, and we handle empty fields (like "f") by defaulting to an empty string.

Running this code will output exactly the structured data you need, matching your expected fields for all three rows.

内容的提问来源于stack exchange,提问作者Ling

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:26:29