You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python网页爬虫代码无法提取li标签内文本,是什么原因?

问题原因

  • 语法使用错误:BeautifulSoup 中find_all()方法仅支持传入标签名、属性筛选规则等参数,不能直接传入CSS选择器完整字符串,要执行CSS选择器查询需要使用select()方法。
  • 内容加载逻辑错误:你要提取的模态框内的列表数据属于JavaScript动态渲染内容,直接通过requests请求得到的静态页面源码中不包含这部分数据,这是你拿到的result长度始终为0的核心原因。

解决方案

方案1:使用Selenium模拟浏览器渲染(适合快速实现)

from bs4 import BeautifulSoup
from selenium import webdriver
from selenium.webdriver.chrome.options import Options
import time

# 配置无头浏览器(无需弹出浏览器窗口)
chrome_options = Options()
chrome_options.add_argument("--headless")
driver = webdriver.Chrome(options=chrome_options)

url = "https://archief.amsterdam/inventarissen/scans/31245/120.3"
driver.get(url)
# 等待页面JS渲染完成,可根据网络情况调整等待时长
time.sleep(3)

soup = BeautifulSoup(driver.page_source, 'html.parser')
# 去掉选择器中的:nth-child(1)即可提取所有li下的file-name内容
result = soup.select('#modal > div > div.content > div > div > ul > li > span.file-name')

for item in result:
    print(item.text.strip())

print(f"共提取到{len(result)}条数据")
driver.quit()

方案2:抓后台接口请求(效率更高)

打开浏览器开发者工具的「网络」面板,筛选XHR/Fetch类请求,找到加载模态框列表数据的后台接口,直接请求接口即可拿到结构化的JSON数据,无需解析HTML。

内容的提问来源于stack exchange,提问作者Fez Talmest

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.28 00:24:05