You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python从HTML中提取指定文本?

用BeautifulSoup提取HTML中的目标文本

1. 先安装必要库

打开终端执行以下命令:

pip install beautifulsoup4 lxml

2. 修改代码实现文本提取

基于你现有的代码,加入BeautifulSoup的解析逻辑,以下是完整示例:

from urllib.request import urlopen
from bs4 import BeautifulSoup

url = input('Please enter the URL to get the HTML from:')

if url:
    response = urlopen(url)
    html = response.read()
    # 用lxml解析器处理HTML
    soup = BeautifulSoup(html, 'lxml')
    
    # 方法1:提取页面所有纯文本(自动过滤标签)
    full_text = soup.get_text(strip=True)
    print("页面所有文本:")
    print(full_text)
    
    # 方法2:精准提取目标文本(需定位对应HTML元素)
    # 假设目标文本在帖子内容对应的div中,可根据实际HTML结构调整属性
    target_element = soup.find('div', class_='userContent')
    if target_element:
        print("\n提取到的目标文本:")
        print(target_element.get_text(strip=True))

关键说明

  • BeautifulSoup(html, 'lxml'):将获取到的HTML字符串转换成可操作的解析对象
  • get_text(strip=True):移除HTML标签提取纯文本,strip=True会自动清理多余的换行和空格
  • find():根据标签名、class/id等属性定位特定元素,如果你能找到目标文本对应的HTML结构,用这个方法能实现精准提取

内容的提问来源于stack exchange,提问作者José A.

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.17 00:05:02