You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python的requests模块获取指定网站国家与首都并生成DataFrame

解决方案:提取世界国家首都并生成DataFrame

用BeautifulSoup解析HTML比正则表达式更可靠——HTML结构的微小变化就会让正则匹配失效,下面是完整实现方案:

1. 安装依赖库

先确保安装所需工具:

pip install requests beautifulsoup4 pandas

2. 完整代码实现

import requests
from bs4 import BeautifulSoup
import pandas as pd

# 发送请求获取页面内容
url = "https://geographyfieldwork.com/WorldCapitalCities.htm"
# 模拟浏览器请求头,避免被拦截
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
}
response = requests.get(url, headers=headers)
response.encoding = 'utf-8'

# 解析HTML页面
soup = BeautifulSoup(response.text, 'html.parser')

# 定位目标表格(页面核心数据都在这个表格里)
table = soup.find('table')

# 初始化数据存储列表
data = []

# 遍历表格行,跳过表头行
for row in table.find_all('tr')[1:]:
    cols = row.find_all('td')
    # 提取国家和首都文本,去除多余空格
    country = cols[0].text.strip()
    capital = cols[1].text.strip()
    # 过滤空行数据
    if country and capital:
        data.append({'country': country, 'capital': capital})

# 创建DataFrame
df = pd.DataFrame(data)

# 打印前5行验证结果
print(df.head())

关键说明

  • 目标页面的国家首都数据封装在HTML表格中,用BeautifulSoup可以精准定位表格结构,避免正则表达式匹配的不确定性。
  • 代码中添加了浏览器请求头,解决部分网站的反爬拦截问题;同时对提取的文本做了去空格、空行过滤处理,保证数据整洁。

内容的提问来源于stack exchange,提问作者Vlad

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 22:25:29