使用Beautiful Soup提取嵌套HTML数据:美国房价爬虫提取问题
Hey Alan, let's get your web scraper pulling that target zip code's housing data correctly. The main issues with your initial code are likely missing request headers (to avoid being blocked), no specific target URL, and incorrect element selection. Here's how to fix it:
Step 1: Add Request Headers & Target URL
Most websites block bare urllib requests because they look like bots. We'll add a User-Agent header to mimic a browser, and specify the actual URL for your target zip code's housing page.
Step 2: Use Precise Element Selectors
You'll need to inspect the target page's HTML to find the right selectors for "Studio" (一室公寓) and "1 Bedroom" (一居室) prices. For example, if the prices are in elements with classes like studio-price and one-bed-price, we'll target those.
Full Modified Code
import urllib.request as urllib2 from bs4 import BeautifulSoup # Replace this with your actual target housing price URL for the zip code target_url = "https://example.com/housing-prices?zip=YOUR_TARGET_ZIP" # Add headers to mimic a browser request headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' } try: # Create a request with headers req = urllib2.Request(target_url, headers=headers) response = urllib2.urlopen(req) html_content = response.read() # Parse the HTML soup = BeautifulSoup(html_content, 'html.parser') # Example selectors - YOU NEED TO UPDATE THESE BASED ON THE ACTUAL PAGE HTML # Inspect the page to find the correct classes/IDs for studio and 1-bed prices studio_price_elem = soup.find('div', class_='studio-price') one_bed_price_elem = soup.find('div', class_='one-bed-price') if studio_price_elem and one_bed_price_elem: studio_price = studio_price_elem.get_text(strip=True) one_bed_price = one_bed_price_elem.get_text(strip=True) print(f"一室公寓 (Studio): {studio_price}") print(f"一居室 (1 Bedroom): {one_bed_price}") else: print("Could not find the price elements. Double-check your selectors!") except Exception as e: print(f"An error occurred: {str(e)}")
Key Notes to Adjust:
- Update the
target_url: Replace the example URL with the actual page you're scraping for your target zip code. - Fix the selectors: Right-click on the price elements in your browser, select "Inspect", and copy the correct class names, IDs, or XPath to use in
soup.find()orsoup.select(). - Handle dynamic content: If the prices load after JavaScript runs (you won't see them in the page source), BeautifulSoup won't catch them. For this case, use tools like
seleniumorrequests-htmlto render the JS before parsing.
Example for Dynamic Content (Using Selenium)
If the page uses JS to load prices, here's a quick alternative snippet:
from selenium import webdriver from selenium.webdriver.common.by import By import time driver = webdriver.Chrome() driver.get(target_url) time.sleep(3) # Wait for JS to load studio_price = driver.find_element(By.CLASS_NAME, 'studio-price').text one_bed_price = driver.find_element(By.CLASS_NAME, 'one-bed-price').text print(f"一室公寓 (Studio): {studio_price}") print(f"一居室 (1 Bedroom): {one_bed_price}") driver.quit()
Make sure you have the Selenium ChromeDriver installed if you go this route.
内容的提问来源于stack exchange,提问作者Alan Tang

