You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何批量抓取网页中ID呈序列变化的span标签内容?(Python优先)

Python批量抓取带规律ID的网页元素数据方案

Hey there! Let's tackle this problem—you need to scrape all those <span> elements with IDs like DataListTicker_lblTicker_0, DataListTicker_lblTicker_1, up to _n. Here are a couple of straightforward Python approaches using popular scraping libraries:

方法1:使用BeautifulSoup + CSS属性选择器(推荐)

这种方法利用CSS属性选择器匹配ID的前缀,完全不用管具体的数字后缀,灵活性拉满。

首先确保你安装了所需的库:

pip install requests beautifulsoup4

代码示例如下:

import requests
from bs4 import BeautifulSoup

# 替换成你的目标网页链接
url = "YOUR_TARGET_URL_HERE"

# 获取网页内容
response = requests.get(url)
response.raise_for_status()  # 请求失败时直接抛出异常

# 解析HTML
soup = BeautifulSoup(response.text, "html.parser")

# 选中所有ID以"DataListTicker_lblTicker_"开头的span元素
target_elements = soup.select('span[id^="DataListTicker_lblTicker_"]')

# 提取所有元素的文本内容
data_list = [element.get_text(strip=True) for element in target_elements]

# 输出结果
print("抓取到的数据:")
for item in data_list:
    print(item)

原理说明:

select('span[id^="DataListTicker_lblTicker_"]') 里的 id^= 语法,专门用来匹配ID属性值以指定字符串开头的元素,完美适配你的ID命名规律,不管有多少个目标元素都能一次性抓全。

方法2:使用正则表达式匹配ID

如果你需要更严格的匹配(比如确保后缀是纯数字),可以结合BeautifulSoup和正则表达式:

import requests
from bs4 import BeautifulSoup
import re

url = "YOUR_TARGET_URL_HERE"

response = requests.get(url)
response.raise_for_status()
soup = BeautifulSoup(response.text, "html.parser")

# 正则规则:匹配ID为DataListTicker_lblTicker_+数字的格式
id_pattern = re.compile(r"DataListTicker_lblTicker_\d+")
target_elements = soup.find_all("span", id=id_pattern)

data_list = [element.get_text(strip=True) for element in target_elements]

print("抓取到的数据:")
for item in data_list:
    print(item)

这种方式会严格筛选出后缀为数字的span元素,避免误抓其他类似ID的无关元素。

关键注意事项

  • 反爬应对:如果requests.get()无法获取正确内容,可能是网站拦截了自动化请求,可以添加请求头模拟浏览器:
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
}
response = requests.get(url, headers=headers)
  • 动态内容处理:如果数据是通过JavaScript动态渲染的(不在初始HTML里),requests就抓不到了,这时可以改用selenium模拟浏览器加载页面后再抓取。

内容的提问来源于stack exchange,提问作者Avinash Singh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 03:31:10