You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中提取HTML文档指定span标签内的文本

Hey there! Let's fix this problem together. The issue with your current code is that find_all('span') grabs every <span> tag in the document, which is why you're getting all span text. To target just the one with id="1", we can narrow down our selection specifically.

Here's how you can do it properly:

Method 1: Use find() with the id parameter

Since HTML id attributes are meant to be unique, using find() (instead of find_all()) will directly return the exact span we want:

from bs4 import BeautifulSoup

with open("10_01.htm") as fp:
    soup = BeautifulSoup(fp, features="html.parser")

# Locate the span with id="1"
target_span = soup.find('span', id='1')

# Check if the span exists to avoid errors
if target_span:
    # Extract the text inside (strip=True removes extra whitespace/newlines)
    print(target_span.get_text(strip=True))
else:
    print("Couldn't find a span with id='1'")

Method 2: Use CSS selectors

Another flexible way is to use a CSS selector, which works great for more complex queries too:

from bs4 import BeautifulSoup

with open("10_01.htm") as fp:
    soup = BeautifulSoup(fp, features="html.parser")

# Use CSS selector to target span#1
target_span = soup.select_one('span#1')

if target_span:
    print(target_span.get_text(strip=True))
else:
    print("Couldn't find a span with id='1'")

A quick note: I used get_text() instead of .string because .string only works if the span has no child elements. If your span has nested tags inside, .string will return None, but get_text() will still extract all the visible text correctly.

内容的提问来源于stack exchange,提问作者Terry

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 17:12:30