如何从aviasales.com解析信息并抓取指定机票的10个最低价
Got it, let's break down how to solve this problem. The core challenge here is that the flight search results don't load right when the page first renders— the platform needs a few seconds to run its backend search and populate the low-price options. Static scraping tools (like just using requests and BeautifulSoup) won't work here because they only grab the initial HTML, not the dynamically loaded data.
Here's a reliable approach using Playwright (a modern dynamic scraping tool that mimics real browser behavior):
Step 1: Set Up Dependencies
First, install Playwright and the required browser:
pip install playwright playwright install chromium
Step 2: Write the Scraping Script
This script will launch a browser, wait for the results to load, extract all visible prices, sort them, and grab the top 10 lowest ones.
from playwright.sync_api import sync_playwright import re def get_top_10_lowest_prices(url): with sync_playwright() as p: # Launch a headless browser (remove headless=False to see the browser window) browser = p.chromium.launch(headless=True) page = browser.new_page() # Navigate to the target URL, wait until network activity calms down page.goto(url, wait_until="networkidle") # Wait for price elements to load (adjust selector if needed) # Inspect the page to confirm the correct CSS selector for prices page.wait_for_selector(".ticket__price-value", timeout=15000) # Extract all price values price_elements = page.query_selector_all(".ticket__price-value") prices = [] for elem in price_elements: price_text = elem.inner_text().strip() # Clean the price string to get a numeric value numeric_price = re.sub(r"[^\d.]", "", price_text) if numeric_price: prices.append(float(numeric_price)) # Sort and grab the top 10 lowest prices sorted_lowest_prices = sorted(prices)[:10] browser.close() return sorted_lowest_prices # Example usage for your target flight target_flight_url = "http://www.aviasales.com/flights/MOW0608BCN1" top_10_prices = get_top_10_lowest_prices(target_flight_url) print("Top 10 lowest flight prices:", top_10_prices)
Key Tips for Reliability
- Wait Smartly: Instead of fixed
time.sleep()delays, usewait_for_selectorto wait until price elements are actually visible. This avoids missing data if the search takes longer than expected. - Selector Tweaks: If
.ticket__price-valuedoesn't work, right-click a price on the page, select "Inspect", and copy the correct CSS selector for the price elements. - Avoid Blocking: To prevent being flagged as a bot:
- Add small delays between requests with
page.wait_for_timeout(1000) - Set a realistic user-agent string via
page.set_extra_http_headers({"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36..."}) - Don't run the script in rapid succession
- Add small delays between requests with
Alternative: Selenium
If you're more familiar with Selenium, the logic is nearly identical—launch a browser, wait for elements to load, extract and process prices. The main difference is the syntax for waiting and selecting elements.
内容的提问来源于stack exchange,提问作者Naumov Alexandr

