如何用Python的BeautifulSoup遍历XML中的表格区块?
Solution to Iterate Through the Table Block with BeautifulSoup
Got it, let's walk through exactly how to parse and iterate through that TV guide table structure using BeautifulSoup. Here's a practical, step-by-step approach:
Step 1: Setup and Initialize BeautifulSoup
First, make sure you've imported BeautifulSoup and have your target HTML/XML content ready (whether scraped from a URL or stored as a string).
from bs4 import BeautifulSoup # Example content matching your structure (expand with full content as needed) content = """ <h2>Fri 4 May</h2> <table cellspacing="0" cellpadding="12"> <tr> <td class="time ">6:00am</td> <td class="other-details "> <a class="prog-link" href="http://www.tvguide.co.uk/m-detail/157702075/137913159/breakfast" id="308829348" > <div class="title" style="border-left:4px solid #CE3D32"> Breakfast </div> <div class="detail"> A round-up of national and interna...</div> </a> </td> </tr> <!-- Add more rows from your actual content here --> </table> """ # Initialize parser: use 'html.parser' for HTML-like content, 'lxml-xml' for strict XML soup = BeautifulSoup(content, 'html.parser')
Step 2: Locate the Target Table
Since the table is tied to a specific date (<h2>Fri 4 May</h2>), we can link the two to ensure we're parsing the correct table (critical if there are multiple tables on the page):
# Find the date heading first date_heading = soup.find('h2', string='Fri 4 May') # Grab the immediately following table target_table = date_heading.find_next('table', {'cellspacing': '0', 'cellpadding': '12'}) # Alternative: If there's only one table matching those attributes, skip the heading step: # target_table = soup.find('table', {'cellspacing': '0', 'cellpadding': '12'})
Step 3: Iterate Through Rows and Extract Data
Loop through each <tr> in the table to pull out time, title, details, and link for each program:
for row in target_table.find_all('tr'): # Extract show time time_cell = row.find('td', class_='time') show_time = time_cell.get_text(strip=True) if time_cell else "No time listed" # Extract program details cell details_cell = row.find('td', class_='other-details') if details_cell: # Get program link prog_link = details_cell.find('a', class_='prog-link') link_url = prog_link['href'] if prog_link else "No link available" # Get program title title_div = prog_link.find('div', class_='title') show_title = title_div.get_text(strip=True) if title_div else "No title listed" # Get program description detail_div = prog_link.find('div', class_='detail') show_description = detail_div.get_text(strip=True) if detail_div else "No description available" # Output or process the data (adjust as needed) print(f"📅 {date_heading.get_text(strip=True)}") print(f"⏰ {show_time}") print(f"🎬 {show_title}") print(f"ℹ️ {show_description}") print(f"🔗 {link_url}") print("---")
Key Notes for Robustness
- Handle missing elements: Use conditional checks (like
if time_cell) to avoidAttributeErrorif a cell or tag is missing in some rows. - Clean text:
get_text(strip=True)removes extra whitespace, newlines, and tabs from the extracted text. - Parser choice: If you're working with strict XML instead of HTML, swap
'html.parser'with'lxml-xml'(you'll need thelxmllibrary installed for this).
内容的提问来源于stack exchange,提问作者gdogg371
相关产品推荐
相关产品推荐

