You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python 3从HTML列表提取信息至pandas数据结构/列表/字典

如何将网页列表提取为Python可用格式?

嘿,我来帮你搞定这个网页列表提取的事儿!针对你提到的带checkbox的国家选择列表,我整理了几种实用的Python方案,直接就能上手:

方案1:用BeautifulSoup解析静态HTML(最适合快速需求)

这是处理静态HTML最常用的方式,步骤简单易上手:

  1. 先安装依赖包:
pip install beautifulsoup4 requests
  1. 解析代码示例:
from bs4 import BeautifulSoup

# 假设你已经拿到了网页的源代码,存在html变量里
html = """<ul class='checklist__list'> <li class=' checklist__item' id='checklist__item--country-111'> <label class='checklist__label ripple-animation'> <input class="checklist__input js-checklist__input idb-on-change" type="checkbox" id="111" name="country" value="China"> China </label></li>"""

soup = BeautifulSoup(html, 'html.parser')

# 提取所有国家选项
country_list = []
for li in soup.select('ul.checklist__list li.checklist__item'):
    checkbox = li.find('input', class_='checklist__input')
    if checkbox:
        country_info = {
            'id': checkbox.get('id'),
            'value': checkbox.get('value'),
            'name': checkbox.next_sibling.strip()  # 获取标签里的国家名称
        }
        country_list.append(country_info)

print(country_list)
# 输出示例: [{'id': '111', 'value': 'China', 'name': 'China'}]

注意:如果你的HTML是从网页爬取的,先用requests.get(url)获取响应内容再传入BeautifulSoup。

方案2:用Scrapy框架(适合大规模/批量爬取)

如果需要爬取整个网站的多个列表,Scrapy会更高效:

  1. 创建Scrapy项目后,编写爬虫代码:
import scrapy

class CountrySpider(scrapy.Spider):
    name = 'country_spider'
    start_urls = ['你的目标网页URL']

    def parse(self, response):
        for li in response.css('ul.checklist__list li.checklist__item'):
            yield {
                'id': li.css('input.checklist__input::attr(id)').get(),
                'value': li.css('input.checklist__input::attr(value)').get(),
                'name': li.css('label.checklist__label::text').get().strip()
            }

运行爬虫后,数据会自动保存为JSON/CSV等格式,直接就能在Python里使用。

方案3:处理动态加载的列表(JS渲染内容)

如果列表是通过JavaScript动态生成的,BeautifulSoup就抓不到了,这时候用Selenium或者Playwright:

以Selenium为例,先安装依赖:

pip install selenium

代码示例:

from selenium import webdriver
from selenium.webdriver.common.by import By

driver = webdriver.Chrome()  # 需要提前下载对应浏览器的驱动
driver.get('你的目标网页URL')

# 等待列表加载完成(用显式等待会更可靠,这里先简化用隐式等待)
driver.implicitly_wait(10)

country_list = []
for li in driver.find_elements(By.CSS_SELECTOR, 'ul.checklist__list li.checklist__item'):
    checkbox = li.find_element(By.CSS_SELECTOR, 'input.checklist__input')
    country_info = {
        'id': checkbox.get_attribute('id'),
        'value': checkbox.get_attribute('value'),
        'name': checkbox.find_element(By.XPATH, 'following-sibling::text()').get_attribute('textContent').strip()
    }
    country_list.append(country_info)

driver.quit()
print(country_list)

以上三种方案覆盖了绝大多数场景,你可以根据自己的需求选最合适的!

内容的提问来源于stack exchange,提问作者Nini

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 10:24:55