You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python提取文本文件中HTML的href链接与aria-label产品名

提取HTML文本中的href链接与产品名称解决方案

嘿,我来帮你搞定这个提取问题!你的脚本目前能定位到目标行,但还没实现精准提取内容。这里有两种靠谱的实现方式,推荐用HTML解析库来处理,避免正则处理HTML的潜在问题:

方法一:使用BeautifulSoup(推荐,更稳定)

BeautifulSoup是专门的HTML/XML解析库,能轻松处理标签属性提取,不用担心HTML格式小变化导致提取失败。

步骤1:安装依赖

先安装BeautifulSoup库:

pip install beautifulsoup4

步骤2:完整代码实现

from bs4 import BeautifulSoup

# 修正文件路径的定义错误
file_path = '/Users/lins/Downloads/pladiv.txt'

with open(file_path, 'r', encoding='utf-8') as file:
    cnt = 1
    for line in file:
        # 只处理包含目标class的行
        if 'pla-unit-single-clickable-target clickable-card' in line:
            print(f'Position : {cnt}')
            cnt += 1
            # 将当前行作为HTML片段解析
            soup = BeautifulSoup(line, 'html.parser')
            a_tag = soup.find('a')
            
            # 提取href链接
            product_url = a_tag.get('href')
            print(f'URL: {product_url}')
            
            # 提取aria-label并解析产品名称
            aria_label = a_tag.get('aria-label')
            if aria_label:
                # 从aria-label中提取产品名称:截取到"for €"之前的部分
                product_name_part = aria_label.split(' for €')[0]
                # 处理带有价格前缀的情况(比如第一个示例的"0,12 € / 1,00 St. ")
                if ' St. ' in product_name_part:
                    # 取最后一个" St. "之后的内容作为产品名
                    product_name = product_name_part.split(' St. ')[-1].strip()
                else:
                    product_name = product_name_part.strip()
                print(f'产品名称: {product_name}')
            print('---')

方法二:使用正则表达式(适合固定格式的文本)

如果你的HTML格式非常固定,也可以用正则直接匹配属性值,代码更简洁:

import re

file_path = '/Users/lins/Downloads/pladiv.txt'

# 匹配href和aria-label的正则表达式
href_pattern = re.compile(r'href="(.*?)"')
aria_label_pattern = re.compile(r'aria-label="(.*?)"')

with open(file_path, 'r', encoding='utf-8') as file:
    cnt = 1
    for line in file:
        if 'pla-unit-single-clickable-target clickable-card' in line:
            print(f'Position : {cnt}')
            cnt += 1
            
            # 提取href链接
            href_match = href_pattern.search(line)
            if href_match:
                product_url = href_match.group(1)
                print(f'URL: {product_url}')
            
            # 提取aria-label并解析产品名称
            aria_label_match = aria_label_pattern.search(line)
            if aria_label_match:
                aria_label = aria_label_match.group(1)
                product_name_part = aria_label.split(' for €')[0]
                if ' St. ' in product_name_part:
                    product_name = product_name_part.split(' St. ')[-1].strip()
                else:
                    product_name = product_name_part.strip()
                print(f'产品名称: {product_name}')
            print('---')

代码说明

  • 两种方法都先定位到包含目标class的行,避免处理无关内容
  • 提取href属性直接获取完整URL,包括带&的转义字符(如果需要转义为&,可以用product_url.replace('&', '&'))
  • 产品名称的解析逻辑适配了你提供的三种aria-label格式:
    • 对于带价格前缀的(比如第一个示例),去掉前面的“X € / X St. ”部分
    • 对于无价格前缀的(比如第二个示例),直接取“for €”之前的内容

内容的提问来源于stack exchange,提问作者Linu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 09:02:36