如何使用BeautifulSoup筛选表格中最后一列含“oui”的行?
Extract Table Rows Where Last Cell Contains "oui" with BeautifulSoup
Hey there! Let's fix this up for you. You're on the right track switching to BeautifulSoup—regular expressions get really messy with HTML structure, so that was a smart call. Here's how to filter those rows properly:
Step-by-Step Solution
Let's adjust your code to target table rows (<tr>) directly, then check the relevant cell in each row for the "oui" value.
import requests from bs4 import BeautifulSoup url = 'www.someurl.com' headers = {"User-Agent":"Mozilla/5.0"} # Fetch and parse the HTML content response = requests.get(url, headers=headers) html_soup = BeautifulSoup(response.content, 'lxml') # Get all table rows (we'll skip the header row later if needed) all_rows = html_soup.find_all('tr') filtered_rows = [] for row in all_rows: # Grab all cells in the current row cells = row.find_all('td') # Skip the header row (it has bold text in cells, no data) if not cells or 'Liste des candidats' in cells[0].get_text(): continue # Get the last cell (the "Elu(e)" column) last_cell = cells[-1] # Check if the cell contains "oui" (strip whitespace to handle extra spaces/ ) if 'oui' in last_cell.get_text(strip=True): filtered_rows.append(row) # Now you can process the filtered rows, e.g., extract data: for row in filtered_rows: cells = row.find_all('td') candidate = cells[0].get_text(strip=True) votes = cells[1].get_text(strip=True) print(f"Elected Candidate: {candidate}, Votes: {votes}")
Key Explanations
- We start with
<tr>rows instead of just<td>elements so we can keep the entire row intact when we find a match. cells[-1]grabs the last cell in each row (your "Elu(e)" column) using Python's negative indexing to pick the final item in the list of cells.get_text(strip=True)cleans up extra whitespace and characters from the cell text, making it easy to check for "oui" reliably.- We add a check to skip the header row, since it doesn't contain candidate data.
Alternative: Target the Specific Cell Class
Since your "Elu(e)" cells have the tdcd class and are centered, you can be even more precise by targeting that cell directly:
for row in all_rows: # Find the exact "Elu(e)" cell by class and alignment elu_cell = row.find('td', class_='tdcd', align='center') # Only check if the cell exists and contains "oui" if elu_cell and 'oui' in elu_cell.get_text(strip=True): filtered_rows.append(row)
This is helpful if the table structure might change slightly—you'll always target the right column regardless of its position.
内容的提问来源于stack exchange,提问作者Alexandre
相关产品推荐
相关产品推荐

