You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从下拉列表(li标签)抓取国家及费率?爬虫求助

抓取Boss Revolution加拿大站国家及费率方案

1. 问题根源

你之前的代码抓取的是页面顶部导航栏的<ul>,并非目标国家下拉列表。目标国家列表要么是静态隐藏在特定容器中,要么是JavaScript动态加载的,需要针对不同情况调整代码。

2. 静态页面抓取国家列表(若页面源代码可见国家)

直接定位正确的国家列表容器,代码如下:

from bs4 import BeautifulSoup
import requests

url = "https://www.bossrevolution.ca/en-ca/services/international-calling"
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"
}

req = requests.get(url, headers=headers)
soup = BeautifulSoup(req.text, "lxml")

# 替换为实际国家列表的选择器(通过F12查看元素确认)
country_list = soup.find("ul", class_="country-dropdown-list")
if country_list:
    for item in country_list.find_all("li"):
        country_name = item.get_text(strip=True)
        country_link = item.find("a")["href"]
        # 补全完整URL
        full_link = f"https://www.bossrevolution.ca{country_link}" if country_link.startswith("/") else country_link
        print(f"国家:{country_name} | 详情页链接:{full_link}")
else:
    print("未找到静态国家列表,需用动态加载方式抓取")

3. 动态加载页面抓取国家列表(若页面源代码无国家)

如果国家列表是点击下拉按钮后JS加载的,用Selenium模拟浏览器操作:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from bs4 import BeautifulSoup

url = "https://www.bossrevolution.ca/en-ca/services/international-calling"
driver = webdriver.Chrome()  # 需安装对应版本ChromeDriver
driver.get(url)

# 等待并点击国家下拉按钮(替换为实际按钮选择器)
dropdown_btn = WebDriverWait(driver, 10).until(
    EC.element_to_be_clickable((By.CLASS_NAME, "country-select-trigger"))
)
dropdown_btn.click()

# 等待国家列表加载完成
WebDriverWait(driver, 10).until(
    EC.presence_of_element_located((By.CLASS_NAME, "country-list-item"))
)

# 解析页面获取国家数据
soup = BeautifulSoup(driver.page_source, "lxml")
country_items = soup.find_all("li", class_="country-list-item")

country_data = []
for item in country_items:
    name = item.get_text(strip=True)
    link = item.find("a")["href"]
    full_link = f"https://www.bossrevolution.ca{link}" if link.startswith("/") else link
    country_data.append({"name": name, "link": full_link})

driver.quit()

# 输出抓取结果
for entry in country_data:
    print(entry)

4. 批量抓取各国费率

拿到国家详情页链接后,遍历抓取费率:

import requests
from bs4 import BeautifulSoup
import time

headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"
}

# 假设country_data已从上述步骤获取
for country in country_data:
    try:
        req = requests.get(country["link"], headers=headers)
        req.raise_for_status()
        soup = BeautifulSoup(req.text, "lxml")
        
        # 替换为实际费率元素的选择器
        rate_info = soup.find("div", class_="call-rate-display").get_text(strip=True)
        print(f"{country['name']}:{rate_info}")
        
        # 避免频繁请求触发反爬
        time.sleep(1)
    except Exception as e:
        print(f"抓取{country['name']}失败:{str(e)}")

关键提示

  • 所有元素选择器(class、id等)需自己通过浏览器F12开发者工具确认,页面元素可能随网站更新变化
  • 频繁请求会触发反爬机制,务必添加请求间隔
  • 使用Selenium时,要确保浏览器驱动版本与本地浏览器版本匹配

内容的提问来源于stack exchange,提问作者Demi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 11:25:36