Python脚本实现网页指定按钮下载XLS及Docker部署需求
Got it, let's walk through exactly how to containerize your Selenium script so it automatically grabs the latest Tampa PD Calls for Service XLS file on a schedule—no manual work needed. Here's a step-by-step guide tailored to your needs:
First, lay out your files like this to keep things organized:
tampa-pd-xls-downloader/ ├── download_xls.py # Your refined Selenium script ├── Dockerfile # For building the Docker image ├── requirements.txt # Python dependencies └── docker-compose.yml # Simplifies scheduling and file persistence
Tweak your existing script to work in a headless Docker environment (no GUI) and ensure it saves files to a directory we'll mount to your local machine. Here's a polished version:
from selenium import webdriver from selenium.webdriver.chrome.options import Options from selenium.webdriver.common.by import By import time # Configure Chrome for headless container execution chrome_options = Options() chrome_options.add_argument("--headless=new") chrome_options.add_argument("--no-sandbox") chrome_options.add_argument("--disable-dev-shm-usage") # Set download directory (we'll link this to your local machine later) download_dir = "/app/downloads" prefs = {"download.default_directory": download_dir} chrome_options.add_experimental_option("prefs", prefs) # Initialize driver driver = webdriver.Chrome(options=chrome_options) try: # Navigate to the target portal driver.get("https://apps.tampagov.net/CallsForService_Webapp/Default.aspx?type=TPD") # Wait for page to fully load (adjust timeout if needed) time.sleep(5) # Locate and click the export button by its ID export_button = driver.find_element(By.ID, "ctl00$MainContent$btndata") export_button.click() # Wait for download to complete (tweak based on typical file size) time.sleep(10) print("Latest XLS file downloaded successfully!") except Exception as e: print(f"Error during download: {str(e)}") finally: driver.quit()
Create this Dockerfile to package Python, Chrome, ChromeDriver, and your script into a single image:
# Use a lightweight official Python base image FROM python:3.11-slim # Install system dependencies for Chrome and Selenium RUN apt-get update && apt-get install -y \ wget \ unzip \ chromium \ chromium-driver \ && rm -rf /var/lib/apt/lists/* # Set working directory inside the container WORKDIR /app # Install Python dependencies COPY requirements.txt . RUN pip install --no-cache-dir selenium # Copy your script into the container COPY download_xls.py . # Create a directory for downloaded files RUN mkdir -p /app/downloads # Default command (we'll override this for scheduled runs) CMD ["python", "download_xls.py"]
Your requirements.txt only needs this line:
selenium>=4.10.0
To make the script run on a regular schedule (e.g., daily at 2 AM), use Docker Compose with a cron service. Here's the docker-compose.yml:
version: '3.8' services: tampa-xls-downloader: build: . volumes: # Mount a local folder to the container's download directory - ./local-downloads:/app/downloads # Cron schedule: runs daily at 2 AM (adjust the timing to your needs) command: > sh -c "echo '0 2 * * * python /app/download_xls.py' > /etc/cron.d/download-cron && chmod 0644 /etc/cron.d/download-cron && cron && tail -f /var/log/cron.log"
This links a local-downloads folder on your machine to the container—so every time the script runs, the latest XLS file lands directly in your local directory.
- First, create the local downloads folder on your machine:
mkdir local-downloads - Build and start the container in detached mode:
docker-compose up -d --build - Check logs to confirm everything's working:
docker-compose logs -f
- Adjust Timings: The
time.sleep()values might need tweaking if the portal loads slower or the file is larger than usual. - Version Compatibility: Make sure
chromiumandchromium-driverversions in the Dockerfile match—Selenium is finicky about this. - Error Resilience: Add more detailed error handling to the script if you want to get alerts if the portal changes (like the button ID being updated).
内容的提问来源于stack exchange,提问作者user9560028

