You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用网络爬虫抓取WikiCFP表格数据并构建pandas DataFrame

你遇到的核心问题是WikiCFP的表格结构特殊:每条会议数据对应两行<tr>标签,第一行存缩写、会议全称,第二行存举办日期、地点、截稿日期,逐行遍历自然没法拿到完整的单条数据。

修改后的可直接运行代码如下:

import requests
from bs4 import BeautifulSoup
import pandas as pd

# 新增请求头避免被网站反爬拦截
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36"
}

df = pd.DataFrame(columns=["abbreviation", "name", "dates", "place", "deadline"])
url = "http://www.wikicfp.com/cfp/call?conference=computer%20science&page=1"
response = requests.get(url, headers=headers)
soup = BeautifulSoup(response.content, "html.parser")
table = soup.find("table", align="center", cellpadding="3", cellspacing="1", width="100%")

# 提取所有行,跳过表头行
all_rows = table.find_all("tr")[1:]
# 每两行作为一组处理,对应一条完整会议信息
for idx in range(0, len(all_rows), 2):
    first_row_cells = all_rows[idx].find_all("td")
    second_row_cells = all_rows[idx+1].find_all("td")
    # 提取文本并去除多余换行、空格
    item = [
        first_row_cells[0].text.strip(),
        first_row_cells[1].text.strip(),
        second_row_cells[0].text.strip(),
        second_row_cells[1].text.strip(),
        second_row_cells[2].text.strip()
    ]
    # 追加到dataframe末尾
    df.loc[len(df)] = item

# 查看前5条结果验证
print(df.head())
# 可按需导出为csv
# df.to_csv("wikicfp_computer_science.csv", index=False, encoding="utf-8-sig")

主要改动说明:

  • 新增请求头,解决直接请求大概率被拦截返回403的问题
  • 调整遍历逻辑,按步长2分组处理,匹配网站的表格结构
  • 每个字段提取后用strip()清理无效的换行、前后空格,保证数据干净
  • 直接将整理好的单条数据插入dataframe,不需要额外转其他中间结构

内容的提问来源于stack exchange,提问作者FafamKurac123

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.28 14:36:01