You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python的lxml解析HTML提取数据,替代大量str.replace?

用lxml解析HTML并搜索特定内容

一、替代大量str.replace的HTML解析方案

直接用lxml.html模块把请求到的HTML转换成可遍历的节点树,通过XPath或CSS选择器精准提取数据,完全不用靠字符串替换来清理内容。

实操示例:

  1. 导入所需库:
import requests
import lxml.html
import pandas as pd
  1. 修改数据抓取函数,加入解析逻辑:
def grab_data(url):
    response = requests.get(url)
    # 把HTML内容解析成节点树
    tree = lxml.html.fromstring(response.content)
    
    # 示例:提取页面中的学校全称(假设在h1标签内)
    school_name = tree.xpath('//h1/text()')[0].strip() if tree.xpath('//h1/text()') else None
    
    # 示例:提取学校联系邮箱(假设在class为"email"的元素内)
    email = tree.xpath('//*[@class="email"]/text()')[0].strip() if tree.xpath('//*[@class="email"]/text()') else None
    
    # 返回结构化的字典数据
    return {"school_name": school_name, "email": email}

# 应用到你的DataFrame
df['parsed_data'] = df['URL'].apply(grab_data)

后续可以用pd.json_normalize(df['parsed_data'])把字典数据展开成单独的列,直接用于CSV导出。

二、无明确结构时搜索特定短语的方法

如果HTML结构混乱,找不到固定标签或类名,试试这几种方式:

1. 用XPath的contains函数匹配文本

快速定位包含目标短语的文本节点:

# 提取所有包含"Admissions"的文本片段
target_texts = tree.xpath('//text()[contains(., "Admissions")]/text()')
# 清理空白字符
cleaned_texts = [text.strip() for text in target_texts if text.strip()]

2. 结合正则遍历文本节点

针对更复杂的匹配规则(比如带数字的费用信息),遍历所有文本节点后用正则筛选:

import re

# 匹配包含"Fees"和英镑金额的文本
pattern = re.compile(r'Fees.*?£\d+', re.IGNORECASE)

all_texts = tree.xpath('//text()')
matching_texts = []
for text in all_texts:
    stripped_text = text.strip()
    if stripped_text and pattern.search(stripped_text):
        matching_texts.append(stripped_text)

3. 先缩小搜索范围再查找

如果能确定目标内容大概在某个区域(比如页面主体区、footer),先定位父节点再搜索,减少无效内容:

# 先定位到页面主体div,再搜索包含"Contact"的文本
content_area = tree.xpath('//div[@id="main-content"]')[0]
contact_texts = content_area.xpath('.//text()[contains(., "Contact")]/text()')

内容的提问来源于stack exchange,提问作者elksie5000

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.09 18:40:38