每日存储延迟列车ID至Set并定时写入CSV的技术问题排查
Let's break down what's causing the mismatch between your morning records and the 23:59 count, then fix the entire workflow step by step.
Key Issues in Your Current Code
Your
delayedSetgets reset every time you rungetDelayedCSV
Right now, you create a newdelayedSet = set()insidegetDelayedCSVevery time the function runs. That means every 10 seconds, you're starting fresh—all the morning records get discarded immediately after processing. You need a persistent set that lives across function calls to store the day's data.No connection between your data fetching and the counter logic
Yourtrains()function fetches data and createstrainList, but it never callsgetDelayedCSVto add those trains to your set. Also,delayedCounteris never triggered anywhere, so the daily write/clear actions never happen.Syntax & formatting bugs in
delayedCounter- The date format string
'%Y-%m-d'is missing a%befored(should be'%Y-%m-%d'). file.write(len(delayedSet))is invalid: you need to convert the integer length to a string, add a newline, and close the parentheses properly.- You're opening the file in
a+mode but not handling the case where the file might not exist (thougha+creates it, adding a header makes the CSV more usable).
- The date format string
Wrong scheduler trigger for daily tasks
Using an interval trigger for your daily 23:59 write and 00:00 clear is unreliable. You should use a cron trigger to run those actions at exact times every day.
Corrected Code Implementation
First, let's set up a persistent set, fix the data pipeline, and configure the scheduler correctly:
1. Persistent Storage & Data Processing
from datetime import datetime from apscheduler.schedulers.background import BackgroundScheduler # Create a global persistent set to store daily delayed train IDs daily_delayed_set = set() def getDelayedCSV(trainList): word = 'Atrasado' # Add matching train IDs to the global persistent set for train in trainList: if word in train['status']: daily_delayed_set.add(train['id']) # Use add() instead of update() for single IDs def delayedCounter(): now = datetime.now() # Fixed date format current_date = now.strftime('%Y-%m-%d') current_time = now.strftime('%H:%M') with open('delayedData.csv', 'a+') as file: # Move file pointer to start to check if file is empty (for header) file.seek(0) has_header = file.readline() != '' if current_time == '23:59': # Write date and count as a CSV row count = len(daily_delayed_set) if not has_header: file.write('date,delayed_count\n') file.write(f'{current_date},{count}\n') print(f"Written {count} delayed trains for {current_date}") elif current_time == '00:00': daily_delayed_set.clear() print("Cleared delayed set for new day")
2. Updated Data Fetching & Scheduler Setup
def trains(): try: data = list(getRequest()) trainList = getTrain(data) # Pass the trainList to update our delayed set getDelayedCSV(trainList) except Exception as e: print(f"Error fetching/processing data: {e}") # Initialize scheduler scheduler = BackgroundScheduler() # 1. Run data fetching every 10 seconds scheduler.add_job( func=trains, trigger='interval', seconds=10, max_instances=100, misfire_grace_time=None ) # 2. Run daily count write at 23:59 scheduler.add_job( func=delayedCounter, trigger='cron', hour=23, minute=59 ) # 3. Run set clear at 00:00 scheduler.add_job( func=delayedCounter, trigger='cron', hour=0, minute=0 ) scheduler.start() # Keep the script running (for standalone execution) try: while True: pass except (KeyboardInterrupt, SystemExit): scheduler.shutdown()
What Changed & Why
- Persistent
daily_delayed_set: By making this a global variable, we retain all matching train IDs throughout the day instead of resetting every 10 seconds. - Fixed
getDelayedCSV: Now it adds IDs to the global set instead of creating a new one each time. We also useadd()instead ofupdate()since we're adding single IDs (more readable). - Corrected
delayedCounter: Fixed syntax errors, added CSV headers (optional but cleaner), and ensured we write valid CSV rows with dates and counts. - Cron triggers for daily tasks: Instead of relying on interval timing, we use cron to run the write/clear actions at exact daily times—this is far more reliable for daily scheduled tasks.
- Connected the pipeline: The
trains()function now callsgetDelayedCSVto process each batch of data, so new delayed trains are added to the set as they're fetched.
This setup will keep track of all delayed trains throughout the day, write the total count to delayedData.csv at 23:59, clear the set at midnight, and repeat the cycle the next day.
内容的提问来源于stack exchange,提问作者Bruno Teixeira

