You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python的BeautifulSoup遍历XML中的表格区块?

Solution to Iterate Through the Table Block with BeautifulSoup

Got it, let's walk through exactly how to parse and iterate through that TV guide table structure using BeautifulSoup. Here's a practical, step-by-step approach:

Step 1: Setup and Initialize BeautifulSoup

First, make sure you've imported BeautifulSoup and have your target HTML/XML content ready (whether scraped from a URL or stored as a string).

from bs4 import BeautifulSoup

# Example content matching your structure (expand with full content as needed)
content = """
<h2>Fri 4 May</h2>
<table cellspacing="0" cellpadding="12"> 
  <tr> 
    <td class="time ">6:00am</td> 
    <td class="other-details "> 
      <a class="prog-link" href="http://www.tvguide.co.uk/m-detail/157702075/137913159/breakfast" id="308829348" > 
        <div class="title" style="border-left:4px solid #CE3D32"> Breakfast </div> 
        <div class="detail"> A round-up of national and interna...</div>
      </a>
    </td>
  </tr>
  <!-- Add more rows from your actual content here -->
</table>
"""

# Initialize parser: use 'html.parser' for HTML-like content, 'lxml-xml' for strict XML
soup = BeautifulSoup(content, 'html.parser')

Step 2: Locate the Target Table

Since the table is tied to a specific date (<h2>Fri 4 May</h2>), we can link the two to ensure we're parsing the correct table (critical if there are multiple tables on the page):

# Find the date heading first
date_heading = soup.find('h2', string='Fri 4 May')
# Grab the immediately following table
target_table = date_heading.find_next('table', {'cellspacing': '0', 'cellpadding': '12'})

# Alternative: If there's only one table matching those attributes, skip the heading step:
# target_table = soup.find('table', {'cellspacing': '0', 'cellpadding': '12'})

Step 3: Iterate Through Rows and Extract Data

Loop through each <tr> in the table to pull out time, title, details, and link for each program:

for row in target_table.find_all('tr'):
    # Extract show time
    time_cell = row.find('td', class_='time')
    show_time = time_cell.get_text(strip=True) if time_cell else "No time listed"
    
    # Extract program details cell
    details_cell = row.find('td', class_='other-details')
    if details_cell:
        # Get program link
        prog_link = details_cell.find('a', class_='prog-link')
        link_url = prog_link['href'] if prog_link else "No link available"
        
        # Get program title
        title_div = prog_link.find('div', class_='title')
        show_title = title_div.get_text(strip=True) if title_div else "No title listed"
        
        # Get program description
        detail_div = prog_link.find('div', class_='detail')
        show_description = detail_div.get_text(strip=True) if detail_div else "No description available"
        
        # Output or process the data (adjust as needed)
        print(f"📅 {date_heading.get_text(strip=True)}")
        print(f"⏰ {show_time}")
        print(f"🎬 {show_title}")
        print(f"ℹ️ {show_description}")
        print(f"🔗 {link_url}")
        print("---")

Key Notes for Robustness

  • Handle missing elements: Use conditional checks (like if time_cell) to avoid AttributeError if a cell or tag is missing in some rows.
  • Clean text: get_text(strip=True) removes extra whitespace, newlines, and tabs from the extracted text.
  • Parser choice: If you're working with strict XML instead of HTML, swap 'html.parser' with 'lxml-xml' (you'll need the lxml library installed for this).

内容的提问来源于stack exchange,提问作者gdogg371

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 10:55:44