You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python列表推导式编辑爬取的HTML列表元素时出现丢失问题

列表推导式异常原因

你的推导式添加了if "back" in i的过滤条件,只会匹配到包含back字段的3个带背景色属性的td元素,最终生成的新列表长度只有3。而你赋值的目标my_list[1:]原本有6个元素的位置,Python会直接将左侧切片的内容整体替换为右侧长度为3的列表,原本已经处理好的前3个数值元素就会被直接删除,最终得到只有4个元素的结果。
此外你现有逻辑还有两个隐性问题:

  • 你使用list作为变量名,会覆盖Python内置的list类型,后续调用内置列表方法会出现异常,建议更换为raw_list这类自定义名称
  • 你之前替换<td>的逻辑对带style属性的td标签无效,能提取到102这类值属于巧合,逻辑不具备通用性
  • i.replace(i, i[-1])的写法完全冗余,直接取i[-1]即可实现相同效果
可行的处理方案

快速适配现有需求的写法

直接统一处理所有需要清洗的元素即可,无需额外过滤:

import html

example_list = ['12.7x55 mm PS12B',
  '&lt;td style=&quot;border-bottom:solid 2px&quot;&gt;102\n&lt;/td&gt;',
  '&lt;td style=&quot;border-bottom:solid 2px&quot;&gt;46\n&lt;/td&gt;',
  '&lt;td style=&quot;border-bottom:solid 2px&quot;&gt;57\n&lt;/td&gt;',
  '&lt;td style=&quot;border-bottom:solid 2px; background-color:#00990080;&quot;&gt;6\n&lt;/td&gt;',
  '&lt;td style=&quot;border-bottom:solid 2px; background-color:#00640080;&quot;&gt;5\n&lt;/td&gt;',
  '&lt;td style=&quot;border-bottom:solid 2px; background-color:#FB9C0E80;&quot;&gt;4\n&lt;/td&gt;']

my_list = [example_list[0]]
for raw_item in example_list[1:]:
    # 解码HTML转义字符,将&lt;这类实体转换为对应符号
    decoded_item = html.unescape(raw_item)
    # 提取td标签包裹的文本,去除换行和空白符
    content = decoded_item.split('>')[1].split('<')[0].strip()
    my_list.append(content)

运行后得到的my_list就是你期望的结果:['12.7x55 mm PS12B', '102', '46', '57', '6', '5', '4']

通用HTML内容清洗方案

如果后续要处理更复杂的HTML结构,推荐用Python内置的HTML解析模块处理,稳定性更高,不受标签属性变化影响:

import html
from html.parser import HTMLParser

class TDContentExtractor(HTMLParser):
    def __init__(self):
        super().__init__()
        self.in_td = False
        self.content = None
    
    def handle_starttag(self, tag, attrs):
        if tag == 'td':
            self.in_td = True
    
    def handle_data(self, data):
        if self.in_td:
            self.content = data.strip()
    
    def handle_endtag(self, tag):
        if tag == 'td':
            self.in_td = False

extractor = TDContentExtractor()
my_list = [example_list[0]]
for raw_item in example_list[1:]:
    extractor.feed(html.unescape(raw_item))
    my_list.append(extractor.content)
    # 重置解析器状态处理下一个元素
    extractor.__init__()

内容的提问来源于stack exchange,提问作者Solebay Sharp

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.27 20:06:07