You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

批量抓取多届国际象棋奥赛数据并生成年度CSV文件技术求助

国际象棋奥赛多年份数据抓取优化需求

我已实现2016年国际象棋奥赛结果数据的抓取代码,现在需要自动完成2012、2014年(每两年一届)的数据抓取。目前已编写部分代码,但不清楚后续完善方向,尤其对文件名生成逻辑存在困惑。

2016年单年份抓取代码

import requests
from bs4 import BeautifulSoup
import pandas as pd

#Imports the HTML into python
url = 'https://www.olimpbase.org/2016/2016te14.html'
requests.get(url)
page = requests.get(url)
print(page)
soup = BeautifulSoup(page.text, 'lxml')

#Subsets the HTML to only get the HTML of our table needed
table = soup.find('table', attrs = {'border': '1'})
print(table)

#Gets all the column headers of our table, but just for the first eleven columns in the webpage
table.find_all('td', class_= 'bog')[1:12]
headers = []
for i in table.find_all('td', class_= 'bog')[1:12]:
    title = i.text.strip()
    headers.append(title)

#Creates a dataframe using the column headers from our table
df = pd.DataFrame(columns = headers)

table.find_all('tr')[3:] #We grab data since the fourth row; the previous ones belong to the headers.

for j in table.find_all('tr')[3:]:
    row_data = j.find_all('td')
    row = [tr.text for tr in row_data][0:11]
    length = len(df)
    df.loc[length] = row

当前编写的多年份抓取代码(未完成)

import requests
from bs4 import BeautifulSoup
import pandas as pd

#Imports the HTML into python
url = 'https://www.olimpbase.org/2016/2016te14.html'
requests.get(url)
page = requests.get(url)
print(page)
soup = BeautifulSoup(page.text, 'lxml')

#Subsets the HTML to only get the HTML of our table needed
table = soup.find('table', attrs = {'border': '1'})
print(table)

#Gets all the column headers of our table
table.find_all('td', class_= 'bog')[1:12]
headers = []
for i in table.find_all('td', class_= 'bog')[1:12]:
    title = i.text.strip()
    headers.append(title)

#Creates a dataframe using the column headers from our table
df = pd.DataFrame(columns = headers)

table.find_all('tr')[3:] #We grab data since the fourth row; the previous ones belong to the headers. 

start_year=2012
i=2
end_year=2016 

def download_chess(start_year):
    url = f'https://www.olimpbase.org/{start_year}/{start_year}te14.html'
    response = requests.get(url)
    soup = BeautifulSoup(page.text, 'lxml')
    for j in table.find_all('tr')[3:]:
        row_data = j.find_all('td')
        row = [tr.text for tr in row_data][0:11]
        length = len(df)
        df.loc[length] = row

while start_year<end_year:
    download_chess(start_year)
    start_year+=i    

download_chess(start_year)

问题分析与解决方案

现有代码的核心问题

  • 全局变量复用错误:download_chess函数未使用当前请求的response,而是复用了全局的2016年page和table变量,导致所有年份抓取的都是2016年数据。
  • 数据混合存储:全局df会将所有年份数据合并,无法区分单年份数据。
  • 缺失文件保存逻辑:没有将抓取的数据导出为文件,也未实现按年份生成文件名的逻辑。

修正后的完整代码

import requests
from bs4 import BeautifulSoup
import pandas as pd

def download_chess(year):
    # 构造对应年份的请求URL
    url = f'https://www.olimpbase.org/{year}/{year}te14.html'
    response = requests.get(url)
    
    # 检查请求是否成功
    if response.status_code != 200:
        print(f"获取{year}年数据失败")
        return
    
    soup = BeautifulSoup(response.text, 'lxml')
    table = soup.find('table', attrs={'border': '1'})
    
    # 检查目标表格是否存在
    if not table:
        print(f"{year}年页面未找到目标表格")
        return
    
    # 提取表头(前11列)
    headers = []
    for header_cell in table.find_all('td', class_='bog')[1:12]:
        headers.append(header_cell.text.strip())
    
    # 初始化当前年份的DataFrame
    year_df = pd.DataFrame(columns=headers)
    
    # 提取表格数据(从第4行开始)
    for row in table.find_all('tr')[3:]:
        row_cells = row.find_all('td')
        # 避免行数据长度不足导致索引越界
        if len(row_cells) >= 11:
            clean_row = [cell.text.strip() for cell in row_cells][:11]
            year_df.loc[len(year_df)] = clean_row
    
    # 生成带年份的文件名并保存为CSV
    filename = f'国际象棋奥赛结果_{year}.csv'
    year_df.to_csv(filename, index=False)
    print(f"{year}年数据已保存至文件: {filename}")

# 遍历2012、2014、2016三个年份
for year in range(2012, 2017, 2):
    download_chess(year)

关键逻辑说明

  • 文件名生成:通过f'国际象棋奥赛结果_{year}.csv'实现按年份命名,{year}会被替换为当前遍历的年份,比如2012年生成国际象棋奥赛结果_2012.csv,保证每个年份数据独立存储。
  • 函数解耦:每个年份的请求、解析、保存流程都在download_chess函数内完成,避免全局变量干扰,逻辑更清晰。
  • 错误防护:添加请求状态码检查和表格存在性验证,避免因网络问题或页面结构变化导致程序异常。

内容的提问来源于stack exchange,提问作者Ivan A. Ramírez Zapata

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.22 19:15:11