使用Python爬取表格时无法获取偶数行数据的技术求助
问题解决:网页爬取仅获取奇数行数据
问题原因
你的代码只抓取了class为datas0的<td>元素,但目标页面的表格是通过交替使用datas0和datas1类名区分奇偶行的,导致偶数行的<td>被完全忽略,最终只能拿到奇数行数据。
解决方案
方法1:同时匹配两个类名的td元素
修改查找td的逻辑,同时包含datas0和datas1类:
import requests from bs4 import BeautifulSoup import pandas as pd # URL to scrape url = 'https://bases.athle.fr/asp.net/liste.aspx?frmbase=resultats&frmmode=1&frmespace=0&frmcompetition=268139&frmposition={page}' # Create an empty list to store the data data = [] # Loop through all pages (assuming there are less than range) for page in range(4): # Send a GET request to the URL response = requests.get(url.format(page=page)) # Create a BeautifulSoup object with the content of the response soup = BeautifulSoup(response.content, 'html.parser') # 同时查找class为datas0和datas1的td元素 td_elements = soup.find_all('td', {'class': ['datas0', 'datas1']}) # Check if there are any results if len(td_elements) == 0: break # Loop through each td element and extract the text for td in td_elements: data.append(td.text.strip()) # Divide the data into columns cols = ['Rank', 'Mark', 'Name', 'Club', 'Department', 'Frm_Ligue','Cat_Sex','Col8','Col9'] # Convert the data into a Pandas dataframe df = pd.DataFrame([data[i:i+len(cols)] for i in range(0, len(data), len(cols))], columns=cols) # Save the dataframe as a CSV file df.to_csv('athle.csv', index=False) print('Data saved to CSV!')
方法2:通过表格行遍历(更健壮)
先定位目标表格,再遍历每一行提取td,这种方法不受类名变化影响:
import requests from bs4 import BeautifulSoup import pandas as pd # URL to scrape url = 'https://bases.athle.fr/asp.net/liste.aspx?frmbase=resultats&frmmode=1&frmespace=0&frmcompetition=268139&frmposition={page}' # Create an empty list to store the data data = [] # Loop through all pages (assuming there are less than range) for page in range(4): # Send a GET request to the URL response = requests.get(url.format(page=page)) # Create a BeautifulSoup object with the content of the response soup = BeautifulSoup(response.content, 'html.parser') # 定位目标表格(页面中包含赛事数据的表格) target_table = soup.find('table', {'class': 'tableau'}) if not target_table: break # 遍历表格的每一行(跳过表头行) for row in target_table.find_all('tr')[1:]: # 提取当前行的所有td文本 row_data = [td.text.strip() for td in row.find_all('td')] # 将行数据添加到总列表 data.extend(row_data) # Divide the data into columns cols = ['Rank', 'Mark', 'Name', 'Club', 'Department', 'Frm_Ligue','Cat_Sex','Col8','Col9'] # Convert the data into a Pandas dataframe df = pd.DataFrame([data[i:i+len(cols)] for i in range(0, len(data), len(cols))], columns=cols) # Save the dataframe as a CSV file df.to_csv('athle.csv', index=False) print('Data saved to CSV!')
说明
- 方法1直接针对类名问题修复,简单快捷;
- 方法2通过表格行遍历,更健壮,即使后续页面调整类名,只要表格结构不变,依然能正常抓取数据。
内容的提问来源于stack exchange,提问作者CarlosFC
相关产品推荐
相关产品推荐

