You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python爬取表格时无法获取偶数行数据的技术求助

问题解决:网页爬取仅获取奇数行数据

问题原因

你的代码只抓取了class为datas0的<td>元素,但目标页面的表格是通过交替使用datas0和datas1类名区分奇偶行的,导致偶数行的<td>被完全忽略,最终只能拿到奇数行数据。

解决方案

方法1:同时匹配两个类名的td元素

修改查找td的逻辑,同时包含datas0和datas1类:

import requests
from bs4 import BeautifulSoup
import pandas as pd

# URL to scrape
url = 'https://bases.athle.fr/asp.net/liste.aspx?frmbase=resultats&frmmode=1&frmespace=0&frmcompetition=268139&frmposition={page}'

# Create an empty list to store the data
data = []

# Loop through all pages (assuming there are less than range)
for page in range(4):
    # Send a GET request to the URL
    response = requests.get(url.format(page=page))

    # Create a BeautifulSoup object with the content of the response
    soup = BeautifulSoup(response.content, 'html.parser')

    # 同时查找class为datas0和datas1的td元素
    td_elements = soup.find_all('td', {'class': ['datas0', 'datas1']})

    # Check if there are any results
    if len(td_elements) == 0:
        break

    # Loop through each td element and extract the text
    for td in td_elements:
        data.append(td.text.strip())

# Divide the data into columns
cols = ['Rank', 'Mark', 'Name', 'Club', 'Department', 'Frm_Ligue','Cat_Sex','Col8','Col9']

# Convert the data into a Pandas dataframe
df = pd.DataFrame([data[i:i+len(cols)] for i in range(0, len(data), len(cols))], columns=cols)

# Save the dataframe as a CSV file
df.to_csv('athle.csv', index=False)

print('Data saved to CSV!')

方法2:通过表格行遍历(更健壮)

先定位目标表格,再遍历每一行提取td,这种方法不受类名变化影响:

import requests
from bs4 import BeautifulSoup
import pandas as pd

# URL to scrape
url = 'https://bases.athle.fr/asp.net/liste.aspx?frmbase=resultats&frmmode=1&frmespace=0&frmcompetition=268139&frmposition={page}'

# Create an empty list to store the data
data = []

# Loop through all pages (assuming there are less than range)
for page in range(4):
    # Send a GET request to the URL
    response = requests.get(url.format(page=page))

    # Create a BeautifulSoup object with the content of the response
    soup = BeautifulSoup(response.content, 'html.parser')

    # 定位目标表格(页面中包含赛事数据的表格)
    target_table = soup.find('table', {'class': 'tableau'})
    if not target_table:
        break

    # 遍历表格的每一行(跳过表头行)
    for row in target_table.find_all('tr')[1:]:
        # 提取当前行的所有td文本
        row_data = [td.text.strip() for td in row.find_all('td')]
        # 将行数据添加到总列表
        data.extend(row_data)

# Divide the data into columns
cols = ['Rank', 'Mark', 'Name', 'Club', 'Department', 'Frm_Ligue','Cat_Sex','Col8','Col9']

# Convert the data into a Pandas dataframe
df = pd.DataFrame([data[i:i+len(cols)] for i in range(0, len(data), len(cols))], columns=cols)

# Save the dataframe as a CSV file
df.to_csv('athle.csv', index=False)

print('Data saved to CSV!')

说明

  • 方法1直接针对类名问题修复,简单快捷;
  • 方法2通过表格行遍历,更健壮,即使后续页面调整类名,只要表格结构不变,依然能正常抓取数据。

内容的提问来源于stack exchange,提问作者CarlosFC

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 14:27:11