You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Beautiful Soup提取嵌套HTML数据:美国房价爬虫提取问题

Fixing Your US Housing Price Scraper with BeautifulSoup4

Hey Alan, let's get your web scraper pulling that target zip code's housing data correctly. The main issues with your initial code are likely missing request headers (to avoid being blocked), no specific target URL, and incorrect element selection. Here's how to fix it:

Step 1: Add Request Headers & Target URL

Most websites block bare urllib requests because they look like bots. We'll add a User-Agent header to mimic a browser, and specify the actual URL for your target zip code's housing page.

Step 2: Use Precise Element Selectors

You'll need to inspect the target page's HTML to find the right selectors for "Studio" (一室公寓) and "1 Bedroom" (一居室) prices. For example, if the prices are in elements with classes like studio-price and one-bed-price, we'll target those.

Full Modified Code

import urllib.request as urllib2
from bs4 import BeautifulSoup

# Replace this with your actual target housing price URL for the zip code
target_url = "https://example.com/housing-prices?zip=YOUR_TARGET_ZIP"

# Add headers to mimic a browser request
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'
}

try:
    # Create a request with headers
    req = urllib2.Request(target_url, headers=headers)
    response = urllib2.urlopen(req)
    html_content = response.read()

    # Parse the HTML
    soup = BeautifulSoup(html_content, 'html.parser')

    # Example selectors - YOU NEED TO UPDATE THESE BASED ON THE ACTUAL PAGE HTML
    # Inspect the page to find the correct classes/IDs for studio and 1-bed prices
    studio_price_elem = soup.find('div', class_='studio-price')
    one_bed_price_elem = soup.find('div', class_='one-bed-price')

    if studio_price_elem and one_bed_price_elem:
        studio_price = studio_price_elem.get_text(strip=True)
        one_bed_price = one_bed_price_elem.get_text(strip=True)
        print(f"一室公寓 (Studio): {studio_price}")
        print(f"一居室 (1 Bedroom): {one_bed_price}")
    else:
        print("Could not find the price elements. Double-check your selectors!")

except Exception as e:
    print(f"An error occurred: {str(e)}")

Key Notes to Adjust:

  • Update the target_url: Replace the example URL with the actual page you're scraping for your target zip code.
  • Fix the selectors: Right-click on the price elements in your browser, select "Inspect", and copy the correct class names, IDs, or XPath to use in soup.find() or soup.select().
  • Handle dynamic content: If the prices load after JavaScript runs (you won't see them in the page source), BeautifulSoup won't catch them. For this case, use tools like selenium or requests-html to render the JS before parsing.

Example for Dynamic Content (Using Selenium)

If the page uses JS to load prices, here's a quick alternative snippet:

from selenium import webdriver
from selenium.webdriver.common.by import By
import time

driver = webdriver.Chrome()
driver.get(target_url)
time.sleep(3)  # Wait for JS to load

studio_price = driver.find_element(By.CLASS_NAME, 'studio-price').text
one_bed_price = driver.find_element(By.CLASS_NAME, 'one-bed-price').text

print(f"一室公寓 (Studio): {studio_price}")
print(f"一居室 (1 Bedroom): {one_bed_price}")

driver.quit()

Make sure you have the Selenium ChromeDriver installed if you go this route.

内容的提问来源于stack exchange,提问作者Alan Tang

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 09:20:59