批量抓取多届国际象棋奥赛数据并生成年度CSV文件技术求助
国际象棋奥赛多年份数据抓取优化需求
我已实现2016年国际象棋奥赛结果数据的抓取代码,现在需要自动完成2012、2014年(每两年一届)的数据抓取。目前已编写部分代码,但不清楚后续完善方向,尤其对文件名生成逻辑存在困惑。
2016年单年份抓取代码
import requests from bs4 import BeautifulSoup import pandas as pd #Imports the HTML into python url = 'https://www.olimpbase.org/2016/2016te14.html' requests.get(url) page = requests.get(url) print(page) soup = BeautifulSoup(page.text, 'lxml') #Subsets the HTML to only get the HTML of our table needed table = soup.find('table', attrs = {'border': '1'}) print(table) #Gets all the column headers of our table, but just for the first eleven columns in the webpage table.find_all('td', class_= 'bog')[1:12] headers = [] for i in table.find_all('td', class_= 'bog')[1:12]: title = i.text.strip() headers.append(title) #Creates a dataframe using the column headers from our table df = pd.DataFrame(columns = headers) table.find_all('tr')[3:] #We grab data since the fourth row; the previous ones belong to the headers. for j in table.find_all('tr')[3:]: row_data = j.find_all('td') row = [tr.text for tr in row_data][0:11] length = len(df) df.loc[length] = row
当前编写的多年份抓取代码(未完成)
import requests from bs4 import BeautifulSoup import pandas as pd #Imports the HTML into python url = 'https://www.olimpbase.org/2016/2016te14.html' requests.get(url) page = requests.get(url) print(page) soup = BeautifulSoup(page.text, 'lxml') #Subsets the HTML to only get the HTML of our table needed table = soup.find('table', attrs = {'border': '1'}) print(table) #Gets all the column headers of our table table.find_all('td', class_= 'bog')[1:12] headers = [] for i in table.find_all('td', class_= 'bog')[1:12]: title = i.text.strip() headers.append(title) #Creates a dataframe using the column headers from our table df = pd.DataFrame(columns = headers) table.find_all('tr')[3:] #We grab data since the fourth row; the previous ones belong to the headers. start_year=2012 i=2 end_year=2016 def download_chess(start_year): url = f'https://www.olimpbase.org/{start_year}/{start_year}te14.html' response = requests.get(url) soup = BeautifulSoup(page.text, 'lxml') for j in table.find_all('tr')[3:]: row_data = j.find_all('td') row = [tr.text for tr in row_data][0:11] length = len(df) df.loc[length] = row while start_year<end_year: download_chess(start_year) start_year+=i download_chess(start_year)
问题分析与解决方案
现有代码的核心问题
- 全局变量复用错误:
download_chess函数未使用当前请求的response,而是复用了全局的2016年page和table变量,导致所有年份抓取的都是2016年数据。 - 数据混合存储:全局
df会将所有年份数据合并,无法区分单年份数据。 - 缺失文件保存逻辑:没有将抓取的数据导出为文件,也未实现按年份生成文件名的逻辑。
修正后的完整代码
import requests from bs4 import BeautifulSoup import pandas as pd def download_chess(year): # 构造对应年份的请求URL url = f'https://www.olimpbase.org/{year}/{year}te14.html' response = requests.get(url) # 检查请求是否成功 if response.status_code != 200: print(f"获取{year}年数据失败") return soup = BeautifulSoup(response.text, 'lxml') table = soup.find('table', attrs={'border': '1'}) # 检查目标表格是否存在 if not table: print(f"{year}年页面未找到目标表格") return # 提取表头(前11列) headers = [] for header_cell in table.find_all('td', class_='bog')[1:12]: headers.append(header_cell.text.strip()) # 初始化当前年份的DataFrame year_df = pd.DataFrame(columns=headers) # 提取表格数据(从第4行开始) for row in table.find_all('tr')[3:]: row_cells = row.find_all('td') # 避免行数据长度不足导致索引越界 if len(row_cells) >= 11: clean_row = [cell.text.strip() for cell in row_cells][:11] year_df.loc[len(year_df)] = clean_row # 生成带年份的文件名并保存为CSV filename = f'国际象棋奥赛结果_{year}.csv' year_df.to_csv(filename, index=False) print(f"{year}年数据已保存至文件: {filename}") # 遍历2012、2014、2016三个年份 for year in range(2012, 2017, 2): download_chess(year)
关键逻辑说明
- 文件名生成:通过
f'国际象棋奥赛结果_{year}.csv'实现按年份命名,{year}会被替换为当前遍历的年份,比如2012年生成国际象棋奥赛结果_2012.csv,保证每个年份数据独立存储。 - 函数解耦:每个年份的请求、解析、保存流程都在
download_chess函数内完成,避免全局变量干扰,逻辑更清晰。 - 错误防护:添加请求状态码检查和表格存在性验证,避免因网络问题或页面结构变化导致程序异常。
内容的提问来源于stack exchange,提问作者Ivan A. Ramírez Zapata
相关产品推荐
相关产品推荐

