You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python网页爬取仅返回最后一个数据点,如何获取全部数据?

爬取所有国家数据的修复方案

你的问题出在循环里每次都用data = [country_name, capital, population]重新赋值,这会把之前的国家数据全部覆盖,最后自然只剩最后一个国家的信息。

修复方法很简单:先初始化一个空列表,每次循环把单个国家的数据追加到列表里,而不是覆盖原有内容。修改后的代码如下:

import requests
from bs4 import BeautifulSoup  # 原代码遗漏了BeautifulSoup的导入,需补上

html_text = requests.get('https://www.scrapethissite.com/pages/simple/').text
soup = BeautifulSoup(html_text, 'lxml')
countrys = soup.find_all('div', class_='col-md-4 country')

# 初始化空列表存储所有国家数据
data = []
for country in countrys:
    country_name = country.find('h3', class_='country-name').text.strip()
    capital = country.find('span', class_='country-capital').text
    population = country.find('span', class_='country-population').text
    # 将当前国家的数据追加到列表中
    data.append([country_name, capital, population])

print(data)

运行修改后的代码,就能得到所有国家的完整数据列表了。

内容的提问来源于stack exchange,提问作者John_Falco

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 20:05:36