You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python+BeautifulSoup获取CSS选择器HTML报KeyError: 'value'如何解决

问题解决方法

报错原因

你触发KeyError: 'value'是因为目标<img>标签没有value属性,你尝试读取不存在的属性就会抛出该错误。

修正方案

你要获取指定元素的完整HTML代码,直接将BeautifulSoup返回的Tag对象转为字符串即可。另外补充了几个常见的优化点避免后续出问题:

  • 打开输出文件时指定utf-8编码,避免中文/特殊字符乱码
  • 给requests加UA请求头,降低被站点反爬拦截的概率
  • BeautifulSoup显式指定解析器,避免运行时警告
  • 加空行跳过逻辑,避免links.txt里的空行导致报错

修改后完整代码

import os
import requests
from bs4 import BeautifulSoup

# 替换为你自己的实际存储路径
workingpath = "./save_dir"

with open("links.txt", "r", encoding="utf-8") as a_file:
    for line in a_file:
        stripped_line = line.strip()
        # 跳过空行
        if not stripped_line:
            continue
        endpoint = stripped_line
        start = stripped_line.find('/tag/') + 5
        end = stripped_line.find('.html', start)
        filename = stripped_line[start:end]
        # 加UA模拟浏览器请求
        resp = requests.get(endpoint, headers={
            "User-Agent": "Mozilla/5.0 (Macintosh; Intel Mac OS X 13_5) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
        })
        # 显式指定解析器
        soup = BeautifulSoup(resp.text, "html.parser")
        completeName = os.path.join(workingpath, f"{filename}.txt")
        with open(completeName, "w", encoding="utf-8") as f_out:
            for inp in soup.select('.card-img-top'):
                # 转字符串即可获取元素完整HTML代码
                f_out.write(f"{str(inp)}\n")

额外说明

如果你只需要提取元素的某个属性,比如图片链接src、替代文本alt,推荐用get方法读取,避免属性不存在时抛出KeyError:

  • 取src:inp.get("src")
  • 取alt:inp.get("alt")

内容的提问来源于stack exchange,提问作者Katherine Elizabeth Kath

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.06 14:39:00