如何使用Python Requests抓取目标数据?纽约市日托数据抓取求助
Fixing NYC Daycare Data Scraping with Requests/Selenium
Hey there! Let's break down why your requests.get and requests.post calls aren't returning the daycare table data, and how to fix it.
Common Issues & Solutions
1. Missing Required Form Data (POST Requests)
Most search forms on websites need specific parameters to trigger the search and return results. When you send a bare GET or empty POST, the server doesn't recognize you're asking for filtered daycare data. Here's how to fix this:
- Open your browser's DevTools (F12), go to the Network tab, then click the site's search button.
- Look for the POST request sent to
SearchAction2.do—check the Form Data section to see all parameters the server expects (likeborough,zipCode,csrfToken, or anactionfield set tosearch). - Use
requests.Session()to maintain cookies between requests (many sites require this to validate your session), then send the POST with all necessary form data.
Example Code with Requests & Session:
import requests from bs4 import BeautifulSoup # Initialize a session to persist cookies across requests session = requests.Session() # First, fetch the initial page to grab cookies and hidden tokens initial_url = "https://a816-healthpsi.nyc.gov/ChildCare/SearchAction2.do" initial_resp = session.get(initial_url) initial_resp.raise_for_status() # Parse the page to extract hidden fields (like CSRF tokens) soup = BeautifulSoup(initial_resp.text, "html.parser") csrf_token = soup.find("input", {"name": "csrfToken"})["value"] # Adjust based on actual field name # Build form data matching what you saw in DevTools form_data = { "csrfToken": csrf_token, "action": "search", "borough": "ALL", # Use specific borough or leave as ALL for all results "zipCode": "", "keywords": "", # Add any other parameters from the Form Data section } # Send the POST request with the complete form data search_resp = session.post(initial_url, data=form_data) search_resp.raise_for_status() # Now check the response for the daycare table print(search_resp.text)
2. Dynamic Content Rendering
If the table loads after the initial page (via JavaScript/AJAX), requests won't capture it because it only fetches static HTML. In this case, you need to simulate a real browser to wait for the JavaScript to load the data.
Example Code with Selenium:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC # Initialize Chrome browser (make sure you have ChromeDriver installed) driver = webdriver.Chrome() driver.get("https://a816-healthpsi.nyc.gov/ChildCare/SearchAction2.do") # Wait for the search button to load, then click it search_btn = WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.ID, "searchButton")) # Adjust button selector as needed ) search_btn.click() # Wait for the daycare table to fully load daycare_table = WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.CSS_SELECTOR, "table#daycareResults")) # Adjust table selector ) # Extract the table's HTML content table_html = daycare_table.get_attribute("outerHTML") print(table_html) # Clean up: close the browser driver.quit()
Key Notes
- Always match the request parameters exactly to what the browser sends—missing even one hidden field can make the server return the initial page instead of results.
- If Selenium feels too heavy, you can try
requests-htmlas a lighter alternative that can render JavaScript without a full browser.
内容的提问来源于stack exchange,提问作者user9463814
相关产品推荐
相关产品推荐

