如何实现定时运行爬虫并将数据追加写入CSV文件?
解决爬虫每日采集天气数据并追加写入CSV的问题
一、修改脚本实现数据追加写入
你的脚本目前用'w'模式打开文件,每次运行都会覆盖原有内容。要实现追加,需要调整文件打开模式,并只在首次运行时写入表头,同时修正原脚本中wind字段的取值错误:
- 引入
os模块判断文件是否存在 - 用
'a'模式打开文件(追加模式) - 修正
wind的选择器,避免数据重复
修改后的代码:
from bs4 import BeautifulSoup from selenium import webdriver from csv import writer import os url = 'https://weather.com/es-US/tiempo/10dias/l/468124beb45cd6817ec6a1da2675f104fede76a75ba7ce08c9ae21606ac208b5' driver = webdriver.Chrome('/Users/thras/Desktop/Chromedriver/chromedriver') driver.get(url) source = driver.page_source soup = BeautifulSoup(source, 'html.parser') # 修正字段提取逻辑,确保每个数据对应正确元素 degrees = soup.find('span', class_="DailyContent--temp--3d4dn").text humidity = soup.find('span', class_="Humidity--value--1YWj_ DailyContent--value--37sk2").text wind = soup.find('span', class_="Wind--windWrapper--3aqXJ DailyContent--value--37sk2").text date = soup.find('span', class_="DailyContent--daypartDate--2A3Wi").text info = [degrees, humidity, wind, date] # 判断文件是否存在,不存在则写入表头 file_exists = os.path.exists('weather.csv') with open('weather.csv', 'a', encoding='utf8', newline='') as f: mywriter = writer(f) if not file_exists: header = ('Degrees', 'Humidity', 'Wind', 'Date') mywriter.writerow(header) mywriter.writerow(info) driver.quit()
二、实现每日自动运行
有两种常用方式实现定时运行:
方法1:使用Python的schedule库
安装schedule库:
pip install schedule
将爬虫逻辑封装成函数,添加定时任务(需保持脚本后台运行):
from bs4 import BeautifulSoup from selenium import webdriver from csv import writer import os import schedule import time def crawl_weather(): url = 'https://weather.com/es-US/tiempo/10dias/l/468124beb45cd6817ec6a1da2675f104fede76a75ba7ce08c9ae21606ac208b5' driver = webdriver.Chrome('/Users/thras/Desktop/Chromedriver/chromedriver') driver.get(url) source = driver.page_source soup = BeautifulSoup(source, 'html.parser') degrees = soup.find('span', class_="DailyContent--temp--3d4dn").text humidity = soup.find('span', class_="Humidity--value--1YWj_ DailyContent--value--37sk2").text wind = soup.find('span', class_="Wind--windWrapper--3aqXJ DailyContent--value--37sk2").text date = soup.find('span', class_="DailyContent--daypartDate--2A3Wi").text info = [degrees, humidity, wind, date] file_exists = os.path.exists('weather.csv') with open('weather.csv', 'a', encoding='utf8', newline='') as f: mywriter = writer(f) if not file_exists: header = ('Degrees', 'Humidity', 'Wind', 'Date') mywriter.writerow(header) mywriter.writerow(info) driver.quit() # 设置每天上午8点运行爬虫 schedule.every().day.at("08:00").do(crawl_weather) while True: schedule.run_pending() time.sleep(60)
方法2:使用系统定时任务(推荐)
无需保持脚本运行,直接利用系统自带功能:
- Mac/Linux:使用crontab
- 终端输入
crontab -e编辑定时任务 - 添加一行(替换为你的Python路径和脚本路径):
表示每天上午8点自动运行脚本。0 8 * * * /usr/bin/python3 /path/to/your/weather_crawler.py
- 终端输入
- Windows:使用任务计划程序
- 搜索打开「任务计划程序」
- 创建基本任务,触发时间设为「每天」,操作选择「启动程序」,指定Python.exe和你的脚本路径。
内容的提问来源于stack exchange,提问作者Alain Procs
相关产品推荐
相关产品推荐

