使用Python 3的BeautifulSoup提取网页标题未获正确结果
Hey there, let's work through why your code is returning "Aviation" instead of the expected "Home - SWF - Stewart International Airport".
First, let's recap your setup and the issue:
- You're using
requestsandBeautifulSoupto fetch and parse the Stewart International Airport homepage - You expect to extract the title declared in the page's HTML:
<title>Home - SWF - Stewart International Airport</title> - But your code is outputting "Aviation" instead
What's probably going wrong
There are two common culprits here:
- Multiple
<title>tags: The raw HTML returned byrequestsmight contain more than one<title>element (for example, from an embedded iframe or a hidden section of the page).soup.titledefaults to grabbing the first one it finds, which happens to be "Aviation". - Dynamic JavaScript rendering: The main page title might be loaded dynamically with JavaScript, which
requestscan't process since it doesn't execute browser-side scripts.
How to debug and fix this
1. Inspect the raw HTML response
First, confirm what title tags are actually present in the data your code is receiving. Add this line right after fetching the response:
print(response.text[:1000]) # Print the first 1000 characters of the raw HTML
This will let you see if the main title is even present in the static HTML.
2. List all title tags in the page
If there are multiple title elements, use find_all to list them all and pick the correct one:
from bs4 import BeautifulSoup import requests url="https://www.swfny.com/" response = requests.get(url) soup = BeautifulSoup(response.text, 'html.parser') # Retrieve all title tags in the document all_titles = soup.find_all('title') for idx, title in enumerate(all_titles): print(f"Title {idx+1}: {title.string}") # If the main title is the second entry, access it like this: # print(all_titles[1].string)
Running this will show you every title tag in the page—chances are the "Home - SWF..." title is further down the list.
3. Handle dynamic content (if needed)
If the main title is loaded via JavaScript, you'll need a tool that can simulate a browser and execute scripts. Here's how to do it with selenium:
from selenium import webdriver from selenium.webdriver.common.by import By # Initialize the browser (make sure ChromeDriver is installed and in your PATH) driver = webdriver.Chrome() driver.get("https://www.swfny.com/") # Extract the rendered page title main_title = driver.find_element(By.TAG_NAME, "title").get_attribute("textContent") print(main_title) # Clean up the browser session driver.quit()
Quick note
In most cases, the multiple title tags scenario is the issue here. Running the find_all('title') code will let you pinpoint exactly which index holds the title you need.
内容的提问来源于stack exchange,提问作者nhrcpt

