如何在Python中提取HTML文档指定span标签内的文本
Hey there! Let's fix this problem together. The issue with your current code is that find_all('span') grabs every <span> tag in the document, which is why you're getting all span text. To target just the one with id="1", we can narrow down our selection specifically.
Here's how you can do it properly:
Method 1: Use find() with the id parameter
Since HTML id attributes are meant to be unique, using find() (instead of find_all()) will directly return the exact span we want:
from bs4 import BeautifulSoup with open("10_01.htm") as fp: soup = BeautifulSoup(fp, features="html.parser") # Locate the span with id="1" target_span = soup.find('span', id='1') # Check if the span exists to avoid errors if target_span: # Extract the text inside (strip=True removes extra whitespace/newlines) print(target_span.get_text(strip=True)) else: print("Couldn't find a span with id='1'")
Method 2: Use CSS selectors
Another flexible way is to use a CSS selector, which works great for more complex queries too:
from bs4 import BeautifulSoup with open("10_01.htm") as fp: soup = BeautifulSoup(fp, features="html.parser") # Use CSS selector to target span#1 target_span = soup.select_one('span#1') if target_span: print(target_span.get_text(strip=True)) else: print("Couldn't find a span with id='1'")
A quick note: I used get_text() instead of .string because .string only works if the span has no child elements. If your span has nested tags inside, .string will return None, but get_text() will still extract all the visible text correctly.
内容的提问来源于stack exchange,提问作者Terry

