如何从含冗余字符的表格行中提取商品标题(Python爬虫)
问题描述
尝试用requests和BeautifulSoup爬取某NFCE网页的表格内容,已找到对应URL并获取到包含表格的HTML元素,但内容里有大量\t、\n等冗余字符,无法正常解析,需要提取每行中类似FEIJAO CARIOCA URBAN的加粗商品标题。
当前代码如下:
from bs4 import BeautifulSoup import requests import re start_url = 'http://nfe.sefaz.go.gov.br/nfeweb/sites/nfce/d/danfeNFCe?p=52240345543915002478650110004179799060499506|2|1|19|118.42|2b4937755a5a76797a522b6435534159784f5859427646356c7a4d3d|1|A763ED574AF1AECE3380D1E3A1EE188E3E95B414' url = "http://nfe.sefaz.go.gov.br/nfeweb/sites/nfce/render/danfeNFCe?chNFe=52240345543915002478650110004179799060499506" headers = { "User-Agent": "Mozilla/5.0 (Macintosh; Intel Mac OS X 10.15; rv:101.0) Gecko/20100101 Firefox/101.0", } with requests.Session() as s: s.headers.update(headers) s.get(start_url) resp = s.get(url) soup = BeautifulSoup(resp.text,"lxml") tabular_content = soup.select_one("script[type='text/javascript']:-soup-contains('new DanfeNFCe')").contents[0] items = re.findall(r"'(<div.*div>)'",tabular_content)[0] print(items)
获取到的表格内容片段:
, GO <\/div> <\/div> <table border=\"0\" align=\"center\" cellpadding=\"0\" cellspacing=\"0\" id=\"tabResult\" data-filter=\"true\"> <tr id=\"Item + 1\"> <td valign=\"top\"><span class=\"txtTit\">FEIJAO CARIOCA URBAN<\/span><span class=\"RCod\"> (Código: 3352048 ) <\/span><br><span class=\"Rqtd\"><strong>Qtde.:<\/strong>2<\/span><span class=\"RUN\"><strong>UN: <\/strong>un<\/span><span class=\"RvlUnit\"><strong>Vl. Unit.:<\/strong> &nbsp; 7,99<\/span><\/td> <td align=\"right\" valign=\"top\" class=\"txtTit noWrap\"> Vl. Total <br><span class=\"valor\">15,98<\/span><\/td> <\/tr> <tr id=\"Item + 2\"> <td valign=\"top\"><span class=\"txtTit\">AZEITONA VDE S CAR<\/span><span class=\"RCod\"> (Código: 5881927 ) <\/span><br><span class=\"Rqtd\"><strong>Qtde.:<\/strong>1<\/span><span class=\"RUN\"><strong>UN: <\/strong>un<\/span><span class=\"RvlUnit\"><strong>Vl. Unit.:<\/strong> &nbsp; 5,59<\/span><\/td>
解决方案
步骤1:处理转义的HTML内容
你当前提取到的items是经过转义的HTML字符串(包含<、>、\/这类转义字符),需要先把这些转义字符还原成正常的HTML标签,再用BeautifulSoup解析。
步骤2:清理冗余字符并提取目标内容
还原后的HTML可能包含大量空白字符(\t、\n、多余空格),可以用BeautifulSoup的get_text()配合字符串清理方法处理,或者直接定位到目标元素提取文本。
修改后的代码如下:
from bs4 import BeautifulSoup import requests import re import html start_url = 'http://nfe.sefaz.go.gov.br/nfeweb/sites/nfce/d/danfeNFCe?p=52240345543915002478650110004179799060499506|2|1|19|118.42|2b4937755a5a76797a522b6435534159784f5859427646356c7a4d3d|1|A763ED574AF1AECE3380D1E3A1EE188E3E95B414' url = "http://nfe.sefaz.go.gov.br/nfeweb/sites/nfce/render/danfeNFCe?chNFe=52240345543915002478650110004179799060499506" headers = { "User-Agent": "Mozilla/5.0 (Macintosh; Intel Mac OS X 10.15; rv:101.0) Gecko/20100101 Firefox/101.0", } with requests.Session() as s: s.headers.update(headers) s.get(start_url) resp = s.get(url) soup = BeautifulSoup(resp.text,"lxml") # 找到包含表格的脚本内容 tabular_content = soup.select_one("script[type='text/javascript']:-soup-contains('new DanfeNFCe')").contents[0] # 提取转义的HTML片段 items = re.findall(r"'(<div.*div>)'",tabular_content)[0] # 1. 还原HTML转义字符 cleaned_html = html.unescape(items) # 2. 将\/替换为/ cleaned_html = cleaned_html.replace('\/', '/') # 3. 用BeautifulSoup解析处理后的HTML item_soup = BeautifulSoup(cleaned_html, "lxml") # 4. 提取所有class为txtTit的span元素文本(排除右侧的Vl. Total) product_titles = [span.get_text(strip=True) for span in item_soup.select("span.txtTit") if not span.parent.has_attr('align')] # 打印结果 for title in product_titles: print(title)
代码说明
html.unescape():将HTML转义字符(如<→<、>→>)还原为正常标签。replace('\/', '/'):处理转义的斜杠,让HTML标签可以被正确解析。select("span.txtTit"):定位到所有商品标题所在的span元素,通过if not span.parent.has_attr('align')排除右侧的“Vl. Total”元素,只保留商品名称。
运行后会输出:
FEIJAO CARIOCA URBAN AZEITONA VDE S CAR
内容的提问来源于stack exchange,提问作者MITHU
相关产品推荐
相关产品推荐

