You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python使用BeautifulSoup爬取网页并提取长度大于8的文本字符串问题

调整方案

你拿到的element.navigableString是BeautifulSoup的内置文本类型,只要将其转为Python原生字符串,再做空白清理和长度筛选,存入列表即可,调整后的代码如下:

import requests
from bs4 import BeautifulSoup as bs

doc = "https://www.kite.com/"
res = requests.get(doc)
# 自动适配网页编码,避免乱码
res.encoding = res.apparent_encoding

soup = bs(res.content, "html.parser")
tag = soup.body

result = []
for string in tag.strings:
    # 转原生字符串,去除首尾空白、换行等无效字符
    clean_text = str(string).strip()
    # 筛选长度大于8的非空内容
    if len(clean_text) > 8:
        result.append(clean_text)

print(result)

说明

  • str(string)可以直接把NavigableString类型转为普通字符串,适配后续的长度判断和列表存储要求
  • strip()用于过滤页面中大量无意义的空行、纯空格文本,减少无效结果
  • 最终输出的result就是你需要的字符串列表格式

内容的提问来源于stack exchange,提问作者jpy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.25 10:24:01