Python BeautifulSoup定位特定元素:提取航班距离数值问题
Hey there! Let's break down why your current code isn't picking up the flight distance value, and get you the "4,866" you need.
The Problem with Your Current Code
Your line dist = soup.find('h2', attrs={'class': 'fa fa-plane'}) is looking for an <h2> tag that has the class fa fa-plane—but if you look at the target HTML:
flight distance = 4,866 miles
That fa fa-plane class belongs to the <i> tag inside the <h2>, not the <h2> itself. So BeautifulSoup can't find the element you want with that selector.
Working Solutions to Extract the Distance
Here are two reliable ways to target the right <h2> and pull out the numerical value:
Method 1: Target the <h2> by its text content
Since the <h2> contains the unique phrase "flight distance", we can use a lambda function to find it:
from bs4 import BeautifulSoup # Assume you've already fetched the page HTML into a variable named 'html' soup = BeautifulSoup(html, 'html.parser') # Find the h2 that contains "flight distance" in its text h2_tag = soup.find('h2', text=lambda t: t and 'flight distance' in t) if h2_tag: # Grab the text inside the <strong> tag distance_value = h2_tag.find('strong').get_text(strip=True) print(distance_value) # Output: 4,866
Method 2: Find the <i> tag first, then navigate to its parent <h2>
Since the <i class="fa fa-plane"> is unique to this <h2>, we can use it as a marker:
from bs4 import BeautifulSoup soup = BeautifulSoup(html, 'html.parser') # Find the i tag with the target class icon_tag = soup.find('i', class_='fa fa-plane') if icon_tag: # Get the parent h2 element h2_tag = icon_tag.parent # Extract the strong tag's text distance_value = h2_tag.find('strong').get_text(strip=True) print(distance_value) # Output: 4,866
Bonus: Clean Up the Numerical Value (If Needed)
If you want to convert "4,866" to an integer for calculations, you can remove the comma:
clean_distance = int(distance_value.replace(',', '')) print(clean_distance) # Output: 4866
内容的提问来源于stack exchange,提问作者Tauseef Hussain

