如何用Python 3从HTML列表提取信息至pandas数据结构/列表/字典
如何将网页列表提取为Python可用格式?
嘿,我来帮你搞定这个网页列表提取的事儿!针对你提到的带checkbox的国家选择列表,我整理了几种实用的Python方案,直接就能上手:
方案1:用BeautifulSoup解析静态HTML(最适合快速需求)
这是处理静态HTML最常用的方式,步骤简单易上手:
- 先安装依赖包:
pip install beautifulsoup4 requests
- 解析代码示例:
from bs4 import BeautifulSoup # 假设你已经拿到了网页的源代码,存在html变量里 html = """<ul class='checklist__list'> <li class=' checklist__item' id='checklist__item--country-111'> <label class='checklist__label ripple-animation'> <input class="checklist__input js-checklist__input idb-on-change" type="checkbox" id="111" name="country" value="China"> China </label></li>""" soup = BeautifulSoup(html, 'html.parser') # 提取所有国家选项 country_list = [] for li in soup.select('ul.checklist__list li.checklist__item'): checkbox = li.find('input', class_='checklist__input') if checkbox: country_info = { 'id': checkbox.get('id'), 'value': checkbox.get('value'), 'name': checkbox.next_sibling.strip() # 获取标签里的国家名称 } country_list.append(country_info) print(country_list) # 输出示例: [{'id': '111', 'value': 'China', 'name': 'China'}]
注意:如果你的HTML是从网页爬取的,先用requests.get(url)获取响应内容再传入BeautifulSoup。
方案2:用Scrapy框架(适合大规模/批量爬取)
如果需要爬取整个网站的多个列表,Scrapy会更高效:
- 创建Scrapy项目后,编写爬虫代码:
import scrapy class CountrySpider(scrapy.Spider): name = 'country_spider' start_urls = ['你的目标网页URL'] def parse(self, response): for li in response.css('ul.checklist__list li.checklist__item'): yield { 'id': li.css('input.checklist__input::attr(id)').get(), 'value': li.css('input.checklist__input::attr(value)').get(), 'name': li.css('label.checklist__label::text').get().strip() }
运行爬虫后,数据会自动保存为JSON/CSV等格式,直接就能在Python里使用。
方案3:处理动态加载的列表(JS渲染内容)
如果列表是通过JavaScript动态生成的,BeautifulSoup就抓不到了,这时候用Selenium或者Playwright:
以Selenium为例,先安装依赖:
pip install selenium
代码示例:
from selenium import webdriver from selenium.webdriver.common.by import By driver = webdriver.Chrome() # 需要提前下载对应浏览器的驱动 driver.get('你的目标网页URL') # 等待列表加载完成(用显式等待会更可靠,这里先简化用隐式等待) driver.implicitly_wait(10) country_list = [] for li in driver.find_elements(By.CSS_SELECTOR, 'ul.checklist__list li.checklist__item'): checkbox = li.find_element(By.CSS_SELECTOR, 'input.checklist__input') country_info = { 'id': checkbox.get_attribute('id'), 'value': checkbox.get_attribute('value'), 'name': checkbox.find_element(By.XPATH, 'following-sibling::text()').get_attribute('textContent').strip() } country_list.append(country_info) driver.quit() print(country_list)
以上三种方案覆盖了绝大多数场景,你可以根据自己的需求选最合适的!
内容的提问来源于stack exchange,提问作者Nini
相关产品推荐
相关产品推荐

