如何用Python自动搜索爬取power of 10网站运动员数据并存入SQLite?
需求可行性与实现方案
这个需求完全可行,下面是具体实现步骤:
一、准备依赖工具
需要用到以下Python库和工具:
selenium:模拟浏览器完成表单填写、点击搜索等交互操作beautifulsoup4:解析网页提取目标数据sqlite3:Python内置的SQLite数据库操作库- 对应浏览器的驱动(比如ChromeDriver,需和浏览器版本匹配)
执行命令安装第三方库:
pip install selenium beautifulsoup4
二、具体实现步骤
1. 模拟浏览器交互,完成搜索操作
用Selenium打开目标网站,定位姓名、俱乐部输入框,填入信息后点击搜索按钮。示例代码如下:
from selenium import webdriver from selenium.webdriver.common.by import By import time # 初始化Chrome浏览器驱动 driver = webdriver.Chrome() driver.get("https://www.thepowerof10.info/") # 替换为实际要搜索的运动员信息 first_name = "John" last_name = "Doe" club = "Example Club" # 定位输入框并填入内容 driver.find_element(By.ID, "firstName").send_keys(first_name) driver.find_element(By.ID, "lastName").send_keys(last_name) driver.find_element(By.ID, "club").send_keys(club) # 点击搜索按钮(需根据页面实际元素调整定位方式) search_btn = driver.find_element(By.XPATH, "//input[@value='Search']") search_btn.click() # 等待搜索结果加载完成 time.sleep(3)
2. 提取搜索结果数据
页面加载完成后,用BeautifulSoup解析页面HTML,提取所需的运动员数据:
from bs4 import BeautifulSoup # 获取当前页面HTML内容 page_source = driver.page_source soup = BeautifulSoup(page_source, "html.parser") # 提取运动员数据(需根据页面实际结构调整选择器) athlete_data = [] # 示例:假设结果在id为athleteResults的表格中 athlete_rows = soup.select("table#athleteResults tr") for row in athlete_rows[1:]: # 跳过表头行 cols = row.find_all("td") if cols: single_data = { "name": cols[0].text.strip(), "club": cols[1].text.strip(), "best_performance": cols[2].text.strip(), # 根据需求添加更多字段 } athlete_data.append(single_data) # 关闭浏览器 driver.quit()
3. 将数据存入SQLite数据库
创建SQLite数据库和表,把提取的数据插入进去:
import sqlite3 # 连接SQLite数据库(不存在则自动创建) conn = sqlite3.connect("athletes.db") cursor = conn.cursor() # 创建数据表(根据实际提取的字段调整结构) cursor.execute(''' CREATE TABLE IF NOT EXISTS athletes ( id INTEGER PRIMARY KEY AUTOINCREMENT, name TEXT NOT NULL, club TEXT, best_performance TEXT ) ''') # 批量插入数据 for data in athlete_data: cursor.execute(''' INSERT INTO athletes (name, club, best_performance) VALUES (?, ?, ?) ''', (data["name"], data["club"], data["best_performance"])) # 提交操作并关闭连接 conn.commit() conn.close()
注意事项
- 页面元素的定位方式(ID、XPATH、CSS选择器等)需要根据网站实际HTML结构调整,建议用浏览器开发者工具查看元素属性
- 遵守网站使用条款,避免过度请求给服务器造成压力
- 若页面加载不稳定,可替换
time.sleep()为WebDriverWait等待特定元素加载,提升代码稳定性
内容的提问来源于stack exchange,提问作者dexta
相关产品推荐
相关产品推荐

