使用Beautiful Soup提取数据返回None的技术求助
Fixing Your Parking Data Extraction with BeautifulSoup
Hey there! Let's break down why your current code isn't pulling the real-time parking spot data you want, and fix it up.
The Core Issue: Wrong Target Structure
First up—you're trying to look for <span class="text"> elements, but the URL you're hitting returns XML data, not HTML. That's why your find_all calls are coming back empty or None—those span tags don't exist in the XML response at all.
If you run print(soup.prettify()), you'll see the actual structure looks something like this (simplified):
<sign_data> <signID id="some-id"> <Display>123</Display> <!-- Other tags --> </signID> <signID id="another-id"> <Display>45</Display> <!-- Other tags --> </signID> </sign_data>
Corrected Code to Extract Display Values
Since we're working with XML, we can directly target the signID and Display tags. Here's the adjusted code:
import requests import bs4 # Fetch the XML data res = requests.get('https://www.jmu.edu/cgi-bin/parking_sign_data.cgi?hash=53616c7465645f5f5c0bbd0eccccb6fe8dd7ed9a0445247e3c7dcb4f91927f7ccc933be780c6e558afb8ebf73620c3e5e3b2c68cd3c138519068eac99d9bf30e1e67ce894deb3a054f95f882da2ea2f0|869835tg89dhkdnbnsv5sg5wg0vmcf4mfcfc2qwm5968unmeh5') res.encoding = 'utf-8' # Ensure proper character encoding soup = bs4.BeautifulSoup(res.text, 'xml') # Extract all Display values from signID tags all_signs = soup.find_all('signID') for sign in all_signs: display_value = sign.find('Display').text print(f"Available Spots: {display_value}") # To target only the Commuter Parking section: # First, check the 'id' attribute of each signID to find the one matching commuter parking # Uncomment the lines below to list all sign IDs: # for sign in all_signs: # print(f"Sign ID: {sign.get('id')}") # Then use that ID to filter: # commuter_sign = soup.find('signID', id='your-commuter-id') # if commuter_sign: # commuter_spots = commuter_sign.find('Display').text # print(f"Commuter Parking Spots Left: {commuter_spots}")
Key Notes
- Always verify the response format first—using
print(soup.prettify())will show you exactly what tags and structure you're working with. - The XML updates every second, so your script will pull the latest real-time data each time it runs.
内容的提问来源于stack exchange,提问作者Ivan Jackson
相关产品推荐
相关产品推荐

