使用Python的requests模块获取指定网站国家与首都并生成DataFrame
解决方案:提取世界国家首都并生成DataFrame
用BeautifulSoup解析HTML比正则表达式更可靠——HTML结构的微小变化就会让正则匹配失效,下面是完整实现方案:
1. 安装依赖库
先确保安装所需工具:
pip install requests beautifulsoup4 pandas
2. 完整代码实现
import requests from bs4 import BeautifulSoup import pandas as pd # 发送请求获取页面内容 url = "https://geographyfieldwork.com/WorldCapitalCities.htm" # 模拟浏览器请求头,避免被拦截 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36' } response = requests.get(url, headers=headers) response.encoding = 'utf-8' # 解析HTML页面 soup = BeautifulSoup(response.text, 'html.parser') # 定位目标表格(页面核心数据都在这个表格里) table = soup.find('table') # 初始化数据存储列表 data = [] # 遍历表格行,跳过表头行 for row in table.find_all('tr')[1:]: cols = row.find_all('td') # 提取国家和首都文本,去除多余空格 country = cols[0].text.strip() capital = cols[1].text.strip() # 过滤空行数据 if country and capital: data.append({'country': country, 'capital': capital}) # 创建DataFrame df = pd.DataFrame(data) # 打印前5行验证结果 print(df.head())
关键说明
- 目标页面的国家首都数据封装在HTML表格中,用
BeautifulSoup可以精准定位表格结构,避免正则表达式匹配的不确定性。 - 代码中添加了浏览器请求头,解决部分网站的反爬拦截问题;同时对提取的文本做了去空格、空行过滤处理,保证数据整洁。
内容的提问来源于stack exchange,提问作者Vlad
相关产品推荐
相关产品推荐

