You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用BeautifulSoup筛选表格中最后一列含“oui”的行?

Extract Table Rows Where Last Cell Contains "oui" with BeautifulSoup

Hey there! Let's fix this up for you. You're on the right track switching to BeautifulSoup—regular expressions get really messy with HTML structure, so that was a smart call. Here's how to filter those rows properly:

Step-by-Step Solution

Let's adjust your code to target table rows (<tr>) directly, then check the relevant cell in each row for the "oui" value.

import requests
from bs4 import BeautifulSoup

url = 'www.someurl.com'
headers = {"User-Agent":"Mozilla/5.0"}

# Fetch and parse the HTML content
response = requests.get(url, headers=headers)
html_soup = BeautifulSoup(response.content, 'lxml')

# Get all table rows (we'll skip the header row later if needed)
all_rows = html_soup.find_all('tr')
filtered_rows = []

for row in all_rows:
    # Grab all cells in the current row
    cells = row.find_all('td')
    # Skip the header row (it has bold text in cells, no data)
    if not cells or 'Liste des candidats' in cells[0].get_text():
        continue
    # Get the last cell (the "Elu(e)" column)
    last_cell = cells[-1]
    # Check if the cell contains "oui" (strip whitespace to handle extra spaces/&nbsp;)
    if 'oui' in last_cell.get_text(strip=True):
        filtered_rows.append(row)

# Now you can process the filtered rows, e.g., extract data:
for row in filtered_rows:
    cells = row.find_all('td')
    candidate = cells[0].get_text(strip=True)
    votes = cells[1].get_text(strip=True)
    print(f"Elected Candidate: {candidate}, Votes: {votes}")

Key Explanations

  • We start with <tr> rows instead of just <td> elements so we can keep the entire row intact when we find a match.
  • cells[-1] grabs the last cell in each row (your "Elu(e)" column) using Python's negative indexing to pick the final item in the list of cells.
  • get_text(strip=True) cleans up extra whitespace and &nbsp; characters from the cell text, making it easy to check for "oui" reliably.
  • We add a check to skip the header row, since it doesn't contain candidate data.

Alternative: Target the Specific Cell Class

Since your "Elu(e)" cells have the tdcd class and are centered, you can be even more precise by targeting that cell directly:

for row in all_rows:
    # Find the exact "Elu(e)" cell by class and alignment
    elu_cell = row.find('td', class_='tdcd', align='center')
    # Only check if the cell exists and contains "oui"
    if elu_cell and 'oui' in elu_cell.get_text(strip=True):
        filtered_rows.append(row)

This is helpful if the table structure might change slightly—you'll always target the right column regardless of its position.

内容的提问来源于stack exchange,提问作者Alexandre

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.11 07:43:21