You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从含冗余字符的表格行中提取商品标题(Python爬虫)

问题描述

尝试用requests和BeautifulSoup爬取某NFCE网页的表格内容,已找到对应URL并获取到包含表格的HTML元素,但内容里有大量\t、\n等冗余字符,无法正常解析,需要提取每行中类似FEIJAO CARIOCA URBAN的加粗商品标题。

当前代码如下:

from bs4 import BeautifulSoup
import requests
import re

start_url = 'http://nfe.sefaz.go.gov.br/nfeweb/sites/nfce/d/danfeNFCe?p=52240345543915002478650110004179799060499506|2|1|19|118.42|2b4937755a5a76797a522b6435534159784f5859427646356c7a4d3d|1|A763ED574AF1AECE3380D1E3A1EE188E3E95B414'
url = "http://nfe.sefaz.go.gov.br/nfeweb/sites/nfce/render/danfeNFCe?chNFe=52240345543915002478650110004179799060499506"

headers = {
    "User-Agent": "Mozilla/5.0 (Macintosh; Intel Mac OS X 10.15; rv:101.0) Gecko/20100101 Firefox/101.0",
}
with requests.Session() as s:
    s.headers.update(headers)
    s.get(start_url)
    resp = s.get(url)
    soup = BeautifulSoup(resp.text,"lxml")
    tabular_content = soup.select_one("script[type='text/javascript']:-soup-contains('new DanfeNFCe')").contents[0]
    items = re.findall(r"'(<div.*div>)'",tabular_content)[0]
    print(items)

获取到的表格内容片段:

,
            							GO
         <\/div>
      <\/div>
      <table border=\"0\" align=\"center\" cellpadding=\"0\" cellspacing=\"0\" id=\"tabResult\" data-filter=\"true\">
         <tr id=\"Item + 1\">
            <td valign=\"top\"><span class=\"txtTit\">FEIJAO CARIOCA URBAN<\/span><span class=\"RCod\">
                  										(Código:
                  										3352048
                  										)
                  									<\/span><br><span class=\"Rqtd\"><strong>Qtde.:<\/strong>2<\/span><span class=\"RUN\"><strong>UN: <\/strong>un<\/span><span class=\"RvlUnit\"><strong>Vl. Unit.:<\/strong>
                  										 
                  										7,99<\/span><\/td>
            <td align=\"right\" valign=\"top\" class=\"txtTit noWrap\">
               									Vl. Total
               									<br><span class=\"valor\">15,98<\/span><\/td>
         <\/tr>
         <tr id=\"Item + 2\">
            <td valign=\"top\"><span class=\"txtTit\">AZEITONA VDE S CAR<\/span><span class=\"RCod\">
                  										(Código:
                  										5881927
                  										)
                  									<\/span><br><span class=\"Rqtd\"><strong>Qtde.:<\/strong>1<\/span><span class=\"RUN\"><strong>UN: <\/strong>un<\/span><span class=\"RvlUnit\"><strong>Vl. Unit.:<\/strong>
                  										 
                  										5,59<\/span><\/td>
解决方案

步骤1:处理转义的HTML内容

你当前提取到的items是经过转义的HTML字符串(包含<、>、\/这类转义字符),需要先把这些转义字符还原成正常的HTML标签,再用BeautifulSoup解析。

步骤2:清理冗余字符并提取目标内容

还原后的HTML可能包含大量空白字符(\t、\n、多余空格),可以用BeautifulSoup的get_text()配合字符串清理方法处理,或者直接定位到目标元素提取文本。

修改后的代码如下:

from bs4 import BeautifulSoup
import requests
import re
import html

start_url = 'http://nfe.sefaz.go.gov.br/nfeweb/sites/nfce/d/danfeNFCe?p=52240345543915002478650110004179799060499506|2|1|19|118.42|2b4937755a5a76797a522b6435534159784f5859427646356c7a4d3d|1|A763ED574AF1AECE3380D1E3A1EE188E3E95B414'
url = "http://nfe.sefaz.go.gov.br/nfeweb/sites/nfce/render/danfeNFCe?chNFe=52240345543915002478650110004179799060499506"

headers = {
    "User-Agent": "Mozilla/5.0 (Macintosh; Intel Mac OS X 10.15; rv:101.0) Gecko/20100101 Firefox/101.0",
}
with requests.Session() as s:
    s.headers.update(headers)
    s.get(start_url)
    resp = s.get(url)
    soup = BeautifulSoup(resp.text,"lxml")
    # 找到包含表格的脚本内容
    tabular_content = soup.select_one("script[type='text/javascript']:-soup-contains('new DanfeNFCe')").contents[0]
    # 提取转义的HTML片段
    items = re.findall(r"'(<div.*div>)'",tabular_content)[0]
    # 1. 还原HTML转义字符
    cleaned_html = html.unescape(items)
    # 2. 将\/替换为/
    cleaned_html = cleaned_html.replace('\/', '/')
    # 3. 用BeautifulSoup解析处理后的HTML
    item_soup = BeautifulSoup(cleaned_html, "lxml")
    # 4. 提取所有class为txtTit的span元素文本(排除右侧的Vl. Total)
    product_titles = [span.get_text(strip=True) for span in item_soup.select("span.txtTit") if not span.parent.has_attr('align')]
    # 打印结果
    for title in product_titles:
        print(title)

代码说明

  • html.unescape():将HTML转义字符(如&lt;→<、&gt;→>)还原为正常标签。
  • replace('\/', '/'):处理转义的斜杠,让HTML标签可以被正确解析。
  • select("span.txtTit"):定位到所有商品标题所在的span元素,通过if not span.parent.has_attr('align')排除右侧的“Vl. Total”元素,只保留商品名称。

运行后会输出:

FEIJAO CARIOCA URBAN
AZEITONA VDE S CAR

内容的提问来源于stack exchange,提问作者MITHU

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.26 14:33:14