You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python3搜索查询后下载CSV文件脚本失效,请求协助修正

Got it, let's break down why your current script isn't pulling the right CSV file and fix it up!

The Problem with Your Current Code

That CSV URL you're hitting directly (http://families.naeyc.org/programs/csv/CSV?attach=list) doesn't know what search results you want. The site generates the CSV based on your active search session—so if you skip the search step and jump straight to the URL, it returns an empty or default file instead of your targeted results.

Solution 1: Simulate Browser Actions with Selenium (Most Reliable)

Since the site relies on frontend interactions (search → wait for results → click download), using Selenium to mimic a real user's browser actions is the most straightforward way to get the correct CSV. Here's how to do it:

First, install Selenium and grab the appropriate browser driver (e.g., ChromeDriver for Chrome):

pip install selenium

Then use this script (adjust selectors to match the site's actual HTML):

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import time

# Configure Chrome to download files to your preferred folder
download_folder = "/your/local/download/path"  # Replace with your actual folder path
chrome_options = webdriver.ChromeOptions()
chrome_options.add_experimental_option(
    "prefs", {"download.default_directory": download_folder}
)

# Launch the browser
driver = webdriver.Chrome(options=chrome_options)

try:
    # Navigate to the main programs page where you perform searches
    driver.get("http://families.naeyc.org/programs")

    # Wait for the search box to load, then enter your keyword and submit
    # Use your browser's dev tools to find the correct selector for the search box
    search_box = WebDriverWait(driver, 10).until(
        EC.presence_of_element_located((By.ID, "search-input"))  # Replace with actual search box selector
    )
    search_box.send_keys("your-search-keyword-here")  # Replace with your search term
    search_box.submit()

    # Wait for search results to load (adjust the selector to match a result element)
    WebDriverWait(driver, 15).until(
        EC.presence_of_element_located((By.CLASS_NAME, "search-result-item"))  # Replace with actual result element selector
    )

    # Find and click the "Download CSV" button
    download_btn = WebDriverWait(driver, 10).until(
        EC.element_to_be_clickable((By.LINK_TEXT, "Download CSV"))  # Replace with button's actual selector if needed
    )
    download_btn.click()

    # Give the file time to download (adjust based on file size)
    time.sleep(5)

finally:
    # Clean up: close the browser
    driver.quit()

Key notes for this script:

  • Use your browser's developer tools (F12) to find the correct CSS/XPath selectors for the search box, results, and download button—these might differ from the placeholders I used.
  • Adjust the wait times (10, 15, 5) based on how fast the site loads for you.

Solution 2: Mimic the Session with Requests (Lightweight Alternative)

If you don't want to use a browser automation tool, you can try mimicking the search session with requests. This requires inspecting the site's network traffic to find how search requests are sent:

import requests

# Create a persistent session to keep cookies and search context
session = requests.Session()

# First, send the search request (you'll need to find the actual search endpoint and parameters)
search_endpoint = "http://families.naeyc.org/programs/search"  # Replace with actual search URL from network dev tools
search_params = {
    "keyword": "your-search-keyword",
    # Add any other required parameters (e.g., filters, page number) you see in the network request
}

# Submit the search (use POST or GET depending on what the site uses)
session.post(search_endpoint, data=search_params)

# Now request the CSV—your session has the search context
csv_url = "http://families.naeyc.org/programs/csv/CSV?attach=list"
response = session.get(csv_url)

# Save the file
with open("targeted_results.csv", "wb") as f:
    f.write(response.content)

This method is lighter but trickier: you'll need to use your browser's network tab to capture exactly what the search request looks like (URL, parameters, headers) to replicate it correctly.

Final Recommendation

Start with the Selenium approach—it's more foolproof for sites that rely on user interactions to generate dynamic content like this CSV.

内容的提问来源于stack exchange,提问作者coder101

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 03:54:42