You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python 3的BeautifulSoup提取网页标题未获正确结果

Fixing Incorrect Title Extraction for Stewart International Airport's Homepage

Hey there, let's work through why your code is returning "Aviation" instead of the expected "Home - SWF - Stewart International Airport".

First, let's recap your setup and the issue:

  • You're using requests and BeautifulSoup to fetch and parse the Stewart International Airport homepage
  • You expect to extract the title declared in the page's HTML: <title>Home - SWF - Stewart International Airport</title>
  • But your code is outputting "Aviation" instead

What's probably going wrong

There are two common culprits here:

  1. Multiple <title> tags: The raw HTML returned by requests might contain more than one <title> element (for example, from an embedded iframe or a hidden section of the page). soup.title defaults to grabbing the first one it finds, which happens to be "Aviation".
  2. Dynamic JavaScript rendering: The main page title might be loaded dynamically with JavaScript, which requests can't process since it doesn't execute browser-side scripts.

How to debug and fix this

1. Inspect the raw HTML response

First, confirm what title tags are actually present in the data your code is receiving. Add this line right after fetching the response:

print(response.text[:1000])  # Print the first 1000 characters of the raw HTML

This will let you see if the main title is even present in the static HTML.

2. List all title tags in the page

If there are multiple title elements, use find_all to list them all and pick the correct one:

from bs4 import BeautifulSoup
import requests
url="https://www.swfny.com/"
response = requests.get(url)
soup = BeautifulSoup(response.text, 'html.parser')

# Retrieve all title tags in the document
all_titles = soup.find_all('title')
for idx, title in enumerate(all_titles):
    print(f"Title {idx+1}: {title.string}")

# If the main title is the second entry, access it like this:
# print(all_titles[1].string)

Running this will show you every title tag in the page—chances are the "Home - SWF..." title is further down the list.

3. Handle dynamic content (if needed)

If the main title is loaded via JavaScript, you'll need a tool that can simulate a browser and execute scripts. Here's how to do it with selenium:

from selenium import webdriver
from selenium.webdriver.common.by import By

# Initialize the browser (make sure ChromeDriver is installed and in your PATH)
driver = webdriver.Chrome()
driver.get("https://www.swfny.com/")

# Extract the rendered page title
main_title = driver.find_element(By.TAG_NAME, "title").get_attribute("textContent")
print(main_title)

# Clean up the browser session
driver.quit()

Quick note

In most cases, the multiple title tags scenario is the issue here. Running the find_all('title') code will let you pinpoint exactly which index holds the title you need.

内容的提问来源于stack exchange,提问作者nhrcpt

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 09:46:42