You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用BeautifulSoup的find_all仅筛选class为even/odd的tr标签?

Fixing Precise Class Filtering with BeautifulSoup

Got it, let's break down why your current code is pulling in those smallrow rows when you only want pure even/odd ones. The key issue here is how BeautifulSoup handles multi-valued class attributes:

When a tag has classes like even smallrow, BeautifulSoup parses that into a list ['even', 'smallrow'] instead of a single string. So your original regex ^(even|odd)$ doesn't check the full class string—it checks individual class values in that list. That means any tag with even or odd in its class list (even if it also has smallrow) gets matched. Same goes for the class_=['even', 'odd'] approach—it matches any tag that has either class, regardless of others.

The Fix: Use a Lambda to Exact-Match Class Lists

To grab only rows that have exactly even or odd (no extra classes), we can use a lambda function in the class_ parameter of find_all. This lets us check if the class attribute is exactly the single value we want:

Updated Code for Pure Even/Odd Rows

Replace your first find_all line with this:

# Only match trs with class exactly "even" or "odd"
for tr in tables[i].find_all('tr', class_=lambda c: c in ['even', 'odd']):

If you want to be extra safe (in case some <tr> tags have no class attribute at all), you can add a check for None:

for tr in tables[i].find_all('tr', class_=lambda c: c is not None and c in ['even', 'odd']):

Refining the Smallrow Filter (Optional)

Your current code for smallrow rows works, but we can make it more robust to handle cases where class order might change (e.g., smallrow even instead of even smallrow). Use another lambda to check that the row has both smallrow and either even or odd:

# Match trs that have smallrow + even/odd, regardless of order
for tr in tables[i].find_all('tr', class_=lambda c: c is not None and ('even' in c or 'odd' in c) and 'smallrow' in c):

Full Corrected Code

Here's your full code with fixes (including the typo request.get → requests.get):

import requests
import pandas as pd
from bs4 import BeautifulSoup as bs

# Fixed typo: request.get → requests.get
page = requests.get('https://regatta.time-team.nl/hollandia/2017/results/003.php')
soup = bs(page.content, 'html.parser')
tables = soup.find_all('table', class_='timeteam')

player_data_even = []
player_data_smallrow = []

for table in tables:
    # Grab only pure even/odd rows
    for tr in table.find_all('tr', class_=lambda c: c in ['even', 'odd']):
        player_row_even = [td.get_text(strip=True) for td in tr.find_all('td')]
        player_data_even.append(player_row_even)
    
    # Grab rows with smallrow + even/odd
    for tr in table.find_all('tr', class_=lambda c: c is not None and ('even' in c or 'odd' in c) and 'smallrow' in c):
        player_row_smallrow = [td.get_text(strip=True) for td in tr.find_all('td')]
        player_data_smallrow.append(player_row_smallrow)

players_even = pd.DataFrame(player_data_even)
players_smallrow = pd.DataFrame(player_data_smallrow)

This will split your rows exactly as you want: pure even/odd rows in one DataFrame, and rows with smallrow plus even/odd in the other.


内容的提问来源于stack exchange,提问作者Jeroen Spaans

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 06:16:58