You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何高效爬取并清洗数据?现有爬取代码是否符合最佳实践?

高效爬取联合国数据及数据清洗优化方案

用户问题描述

我想了解如何高效爬取并清洗数据。目前已使用Python结合requests和BeautifulSoup爬取了联合国网站的数据,但在数据清洗环节遇到困难。以下是我的爬取代码,请问该代码是否符合最佳实践?

import requests
from bs4 import BeautifulSoup
import json

all_countries_links=[]
countries= []
all_data=[]
data_dict={}
data_value=[]
page1 = requests.get(f"https://data.un.org/")

def main(page):
    source = page.content
    soup = BeautifulSoup(source,'lxml')
    all_page = soup.find("div",{"class","CountryList"}).find_all('a',href=True)
    for link in all_page:
        all_countries_links.append(link['href'])
        countries.append(link.text.strip())

def scrape_country(all_countries_links,countries):
     for country in all_countries_links[:2]:
        page2 = requests.get(f"https://data.un.org/{country}") 
        source = page2.content
        soup = BeautifulSoup(source,'lxml')
        all_page= soup.find('ul',{'class','pure-menu-list'})
        tables = all_page.contents
        for table in tables:
            line = table.text.strip()
            all_data.append(line)
main(page1)
scrape_country(all_countries_links,countries)
file_path = "data.json"
with open(file_path, 'w') as f:
    json.dump(all_data, f, indent=4) 
print(f"Data saved to {file_path}")

爬取后的数据示例如下:

[
    "",
    "General Information\n\nRegion\u00a0\n\u00a0\nSouthern Asia\nPopulation\u00a0(000, 2021)\n\u00a0\n39 835a\nPop. density\u00a0(per km2, 2021)\n\u00a0\n61a\nCapital city\u00a0\n\u00a0\nKabul\nCapital city pop.\u00a0(000, 2021)\n\u00a0\n4 114.0b\nUN membership date\u00a0\n\u00a0\n19-Nov-46\nSurface area\u00a0(km2)\n\u00a0\n652 864b\nSex ratio\u00a0(m per 100 f)\n\u00a0\n105.3a\nNational currency\u00a0\n\u00a0\nAfghani (AFN)\nExchange rate\u00a0(per US$)\n\u00a0\n77.1c",   
]

我尝试用以下代码清洗数据,但希望找到更优的方法:

cleaned_data =[]

# for line in cleaned_data:
#     print(line.split('\n'))
# new_data = [line for line in all_data.split()]

for line in all_data[:1]:
    for line2 in line.split():
        if line2 not in ["General","Information","Economic"," indicators","Social"," indicators"]:
            cleaned_data.append(line2)

一、爬取代码的最佳实践优化

你的爬取代码实现了基本功能,但有不少可以优化的点,以下是具体建议:

  • 避免全局变量:当前代码依赖多个全局变量传递数据,易引发变量污染,建议将变量封装到函数内部,通过返回值传递结果。
  • 修复元素匹配语法:find方法中使用{"class","CountryList"}是集合语法,应改为字典{"class": "CountryList"},否则无法正确匹配元素。
  • 添加异常处理:网络请求可能出现超时、连接失败等问题,需捕获requests.exceptions.RequestException类异常,避免程序直接崩溃。
  • 设置请求头:模拟浏览器添加User-Agent,降低被反爬拦截的概率。
  • 限制请求频率:添加time.sleep()延迟,避免短时间内大量请求触发服务器封禁。
  • 结构化爬取:爬取时直接按国家维度组织数据为字典结构,减少后续清洗的复杂度。

优化后的爬取代码示例:

import requests
from bs4 import BeautifulSoup
import json
import time

def get_country_links(base_url):
    headers = {
        "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"
    }
    try:
        response = requests.get(base_url, headers=headers, timeout=10)
        response.raise_for_status()
        soup = BeautifulSoup(response.content, 'lxml')
        country_list = soup.find("div", {"class": "CountryList"})
        if not country_list:
            print("未找到国家列表元素")
            return [], []
        country_links, country_names = [], []
        for a_tag in country_list.find_all('a', href=True):
            country_links.append(a_tag['href'])
            country_names.append(a_tag.text.strip())
        return country_links, country_names
    except requests.exceptions.RequestException as e:
        print(f"获取国家列表失败: {e}")
        return [], []

def scrape_country_data(country_links, country_names, base_url):
    headers = {
        "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"
    }
    country_data = {}
    for idx, link in enumerate(country_links[:2]):
        country_name = country_names[idx]
        country_data[country_name] = {}
        full_url = f"{base_url}/{link}"
        try:
            response = requests.get(full_url, headers=headers, timeout=10)
            response.raise_for_status()
            soup = BeautifulSoup(response.content, 'lxml')
            menu_list = soup.find('ul', {"class": "pure-menu-list"})
            if not menu_list:
                print(f"{country_name}页面未找到数据列表")
                continue
            for li_tag in menu_list.find_all('li'):
                text = li_tag.text.strip()
                if not text:
                    continue
                parts = [p.strip() for p in text.replace('\u00a0', ' ').split('\n') if p.strip()]
                for i in range(0, len(parts), 2):
                    if i+1 < len(parts):
                        country_data[country_name][parts[i]] = parts[i+1]
            time.sleep(1)
        except requests.exceptions.RequestException as e:
            print(f"爬取{country_name}失败: {e}")
            continue
    return country_data

if __name__ == "__main__":
    base_url = "https://data.un.org"
    country_links, country_names = get_country_links(base_url)
    if country_links:
        result = scrape_country_data(country_links, country_names, base_url)
        with open("structured_country_data.json", 'w', encoding='utf-8') as f:
            json.dump(result, f, indent=4, ensure_ascii=False)
        print("结构化数据已保存")

二、数据清洗的更优方案

原始数据的核心特征是“指标名+指标值”成对出现,只是被换行和非断空格干扰。相比简单的字符串拆分,更高效的清洗方式是结构化提取,并做数据标准化处理:

  1. 替换非断空格,过滤无效空行;
  2. 按成对规则提取指标与值;
  3. 可选:去除指标值末尾的标注(如a/b)、转换数值类型。

针对你现有原始数据的清洗代码示例:

import json

# 读取原始数据
with open("data.json", 'r', encoding='utf-8') as f:
    all_data = json.load(f)

cleaned_result = []
for item in all_data:
    if not item.strip():
        continue
    # 替换非断空格,拆分并过滤空内容
    parts = [p.strip() for p in item.replace('\u00a0', ' ').split('\n') if p.strip()]
    # 跳过标题行
    if parts[0] in ["General Information", "Economic indicators", "Social indicators"]:
        parts = parts[1:]
    # 成对提取指标和值
    country_info = {}
    for i in range(0, len(parts), 2):
        if i+1 < len(parts):
            indicator = parts[i]
            value = parts[i+1].rstrip('abc')  # 去除末尾标注
            # 尝试转换数值类型
            try:
                value = float(value.replace(' ', ''))
            except ValueError:
                pass
            country_info[indicator] = value
    cleaned_result.append(country_info)

# 保存清洗后的数据
with open("cleaned_country_data.json", 'w', encoding='utf-8') as f:
    json.dump(cleaned_result, f, indent=4, ensure_ascii=False)
print("数据清洗完成,已保存到cleaned_country_data.json")

内容的提问来源于stack exchange,提问作者basel nabil

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.21 10:13:11