You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python使用schedule定时抓取网页数据写入CSV遇阻求助

Fixing Your Hourly Web Scraping Schedule in Python

Hey there! It looks like you're off to a great start with your Bitcoin Cash scraper—manual runs working is half the battle! Let's tackle the scheduling issue and get that script auto-running every hour smoothly.

Key Issues in Your Current Code

  • Overlapping Scheduled Tasks: You've set up multiple repeating jobs (every minute, every hour, daily at 10:30, etc.). This means your job() function will run way more frequently than you want (e.g., once per minute AND once per hour). For hourly scraping, you only need to keep the hourly schedule rule.
  • Path Escaping Problem: The Windows file path 'C:\Users\NIK\.spyder2\PythonScripts' uses single backslashes, which are escape characters in Python. This can cause the script to fail when trying to change directories.
  • Missing CSV File Extension: You're opening 'BCH_kurs' without a .csv suffix—while this works, it's not standard and might make it harder to open the file with spreadsheet tools later.
  • Lack of Error Handling: If the website is down, the element you're scraping moves, or your internet cuts out, the script will crash and stop running entirely. Adding basic error handling will make it more robust.

Corrected Code

Here's the adjusted version with fixes for all the above issues:

# For Python 3, use urllib.request instead of urllib2
# import urllib.request
import urllib2
from bs4 import BeautifulSoup
import csv
from datetime import datetime
import os
import schedule
import time

def job():
    try:
        # Target URL
        url = 'https://coinmarketcap.com/currencies/bitcoin-cash/'
        # For Python 3, replace with page = urllib.request.urlopen(url)
        page = urllib2.urlopen(url)
        soup = BeautifulSoup(page, 'html.parser')
        
        # Extract Bitcoin Cash value
        name_box = soup.find('span', attrs={'class': 'text-large2'})
        if not name_box:
            print(f"{datetime.now()}: Could not find the price element")
            return
        
        bch_value = float(name_box.text.strip())
        
        # Use raw string for Windows path to avoid escape issues
        os.chdir(r'C:\Users\NIK\.spyder2\PythonScripts')
        # Open CSV file with .csv extension
        with open('BCH_kurs.csv', 'a', newline='') as csv_file:
            writer = csv.writer(csv_file)
            # Write value and timestamp
            writer.writerow([bch_value, datetime.now()])
            
        print(f"{datetime.now()}: Successfully scraped and saved BCH value: {bch_value}")
    
    except Exception as e:
        print(f"{datetime.now()}: Error occurred: {str(e)}")

# Keep ONLY the hourly schedule (remove all other schedule rules)
schedule.every().hour.do(job)

# Run the scheduler loop
while True:
    schedule.run_pending()
    time.sleep(1)

What Changed?

  1. Simplified Scheduling: Removed all extra schedule rules except schedule.every().hour.do(job) to ensure the script runs exactly once per hour.
  2. Fixed File Path: Used a raw string (r'path') for the Windows directory to handle backslashes correctly.
  3. Added CSV Extension: Renamed the output file to BCH_kurs.csv for compatibility with spreadsheet software.
  4. Added Error Handling: Wrapped the entire scraping/writing logic in a try-except block. This catches any errors (like network issues, missing elements) and prints a message instead of crashing the script.
  5. Added Debug Prints: Included status messages so you can see when the script runs successfully or encounters issues.
  6. Python 3 Compatibility Note: Added comments explaining how to switch to urllib.request if you're using Python 3 (since urllib2 is deprecated in Python 3).

Quick Tips

  • If you want the script to run in the background (even when you close your terminal), you can run it as a background process (on Windows, use pythonw.exe instead of python.exe).
  • Test the hourly schedule first by changing it to schedule.every(5).minutes.do(job) to verify it works without waiting an hour.

内容的提问来源于stack exchange,提问作者Nicklas Koldkjær

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 07:58:54