You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup爬取含'is'的<li>标签结果不全问题排查

问题原因

你用string=re.compile("is")筛选<li>标签时,BeautifulSoup的string参数只会匹配标签自身的直接文本内容,不会递归包含子标签里的文本。后面几个目标<li>里嵌套了<i>子标签,它们的string属性无法捕获到包含"is"的完整文本,所以被漏掉了。另外原代码里attrs={'class', 'fun-facts'}写法错误,应该用字典格式{'class': 'fun-facts'}才能正确匹配class属性。

解决方案

换个思路:先获取所有<li>标签,再检查每个标签的完整文本内容(包含子标签文本)是否包含"is",最后提取干净的文本存入列表。

修正后的代码

from bs4 import BeautifulSoup

# 假设webpage是已解析好的BeautifulSoup对象
fun_facts = webpage.find('ul', attrs={'class': 'fun-facts'})
# 列表推导式筛选并提取文本
fun_facts_with_is = [
    li.get_text(strip=True) 
    for li in fun_facts.find_all('li') 
    if 'is' in li.get_text()
]

print(fun_facts_with_is)

代码说明

  • get_text()会将标签及其所有子标签的文本合并成一个字符串,strip=True可以去除文本首尾的空白字符;
  • 遍历所有<li>标签后,通过'is' in li.get_text()判断是否符合条件,避免了string参数的局限性。

运行结果

['Middle name is Ronald',
 'Dunkin Donuts coffee is better than Starbucks',
 "A favorite book series of mine is Ender's Game",
 'Current video game of choice is Rocket League',
 "The band that I've seen the most times live is the Zac Brown Band"]

内容的提问来源于stack exchange,提问作者Josué Lobo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.24 05:22:44