You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python脚本中提取HTML中div下span标签内的3:40文本?

如何用Python提取HTML中的「3:40」文本

嘿,我来帮你搞定这个问题!要从HTML里提取「3:40」这个文本,得先分两种情况:这个时间是直接写在HTML的文本节点里,还是藏在你给的示例图片里?我分别给你讲对应的解决方法:

情况1:「3:40」是HTML文本节点里的内容

这种情况用BeautifulSoup(Python最常用的HTML解析库)就能轻松搞定,步骤如下:

  1. 先安装必要的库:
pip install beautifulsoup4 requests  # requests是用来爬网页的,本地HTML的话可以只装beautifulsoup4
  1. 编写提取代码:
    如果是你已经有了本地的HTML字符串,代码可以这么写:
from bs4 import BeautifulSoup

# 把这里替换成你实际的HTML内容
html_content = '''
<p>视频时长:<span>3:40</span></p>
<p><a href="https://i.sstatic.net/dsQfa.png" rel="nofollow noreferrer">示例元素</a></p>
'''

# 初始化解析器
soup = BeautifulSoup(html_content, 'html.parser')

# 方法1:直接搜索包含「3:40」的文本节点
target_text = soup.find(string=lambda text: text and '3:40' in text.strip())
if target_text:
    extracted_time = target_text.strip()
    print(f"提取到的时间:{extracted_time}")  # 输出:3:40

# 方法2:如果你知道「3:40」所在的标签/类名,直接定位更高效
# 比如假设它在class为duration的span标签里:
# target_element = soup.find('span', class_='duration')
# if target_element:
#     extracted_time = target_element.text.strip()
#     print(f"提取到的时间:{extracted_time}")

如果「3:40」是在某个网页里,你需要先爬取网页内容再解析:

import requests
from bs4 import BeautifulSoup

# 替换成目标网页的URL
url = "你的目标网页地址"
response = requests.get(url)
response.encoding = 'utf-8'  # 确保编码正确,避免乱码

soup = BeautifulSoup(response.text, 'html.parser')

# 同样用上面的方法提取
target_text = soup.find(string=lambda text: text and '3:40' in text.strip())
if target_text:
    print(f"提取到的时间:{target_text.strip()}")

情况2:「3:40」是示例图片里的文字

如果这个时间是在你给的那张图片里(HTML只是引用了图片),那HTML解析工具就没用了,得用**OCR(光学字符识别)**工具,比如pytesseract,步骤如下:

  1. 安装必要的库和OCR引擎:
pip install pytesseract pillow

另外,你还需要安装Tesseract OCR引擎:

  • Windows:从官方渠道下载安装包,记得把安装路径加到系统环境变量里
  • Linux:用sudo apt-get install tesseract-ocr安装
  • macOS:用brew install tesseract安装
  1. 编写OCR提取代码:
from PIL import Image
import pytesseract
import requests
from io import BytesIO

# 图片的URL就是你给的那个链接
image_url = "https://i.sstatic.net/dsQfa.png"
# 获取图片内容
response = requests.get(image_url)
img = Image.open(BytesIO(response.content))

# 用OCR提取图片里的所有文本
full_text = pytesseract.image_to_string(img)

# 从提取的文本里筛选出包含「3:40」的内容
extracted_time = [line.strip() for line in full_text.split('\n') if '3:40' in line.strip()]
if extracted_time:
    print(f"提取到的时间:{extracted_time[0]}")

这样不管是HTML里的文本还是图片里的文字,你都能顺利提取到「3:40」啦!

内容的提问来源于stack exchange,提问作者Jacob Ortman

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 08:48:12