如何高效爬取并清洗数据?现有爬取代码是否符合最佳实践?
高效爬取联合国数据及数据清洗优化方案
用户问题描述
我想了解如何高效爬取并清洗数据。目前已使用Python结合requests和BeautifulSoup爬取了联合国网站的数据,但在数据清洗环节遇到困难。以下是我的爬取代码,请问该代码是否符合最佳实践?
import requests from bs4 import BeautifulSoup import json all_countries_links=[] countries= [] all_data=[] data_dict={} data_value=[] page1 = requests.get(f"https://data.un.org/") def main(page): source = page.content soup = BeautifulSoup(source,'lxml') all_page = soup.find("div",{"class","CountryList"}).find_all('a',href=True) for link in all_page: all_countries_links.append(link['href']) countries.append(link.text.strip()) def scrape_country(all_countries_links,countries): for country in all_countries_links[:2]: page2 = requests.get(f"https://data.un.org/{country}") source = page2.content soup = BeautifulSoup(source,'lxml') all_page= soup.find('ul',{'class','pure-menu-list'}) tables = all_page.contents for table in tables: line = table.text.strip() all_data.append(line) main(page1) scrape_country(all_countries_links,countries) file_path = "data.json" with open(file_path, 'w') as f: json.dump(all_data, f, indent=4) print(f"Data saved to {file_path}")
爬取后的数据示例如下:
[ "", "General Information\n\nRegion\u00a0\n\u00a0\nSouthern Asia\nPopulation\u00a0(000, 2021)\n\u00a0\n39 835a\nPop. density\u00a0(per km2, 2021)\n\u00a0\n61a\nCapital city\u00a0\n\u00a0\nKabul\nCapital city pop.\u00a0(000, 2021)\n\u00a0\n4 114.0b\nUN membership date\u00a0\n\u00a0\n19-Nov-46\nSurface area\u00a0(km2)\n\u00a0\n652 864b\nSex ratio\u00a0(m per 100 f)\n\u00a0\n105.3a\nNational currency\u00a0\n\u00a0\nAfghani (AFN)\nExchange rate\u00a0(per US$)\n\u00a0\n77.1c", ]
我尝试用以下代码清洗数据,但希望找到更优的方法:
cleaned_data =[] # for line in cleaned_data: # print(line.split('\n')) # new_data = [line for line in all_data.split()] for line in all_data[:1]: for line2 in line.split(): if line2 not in ["General","Information","Economic"," indicators","Social"," indicators"]: cleaned_data.append(line2)
一、爬取代码的最佳实践优化
你的爬取代码实现了基本功能,但有不少可以优化的点,以下是具体建议:
- 避免全局变量:当前代码依赖多个全局变量传递数据,易引发变量污染,建议将变量封装到函数内部,通过返回值传递结果。
- 修复元素匹配语法:
find方法中使用{"class","CountryList"}是集合语法,应改为字典{"class": "CountryList"},否则无法正确匹配元素。 - 添加异常处理:网络请求可能出现超时、连接失败等问题,需捕获
requests.exceptions.RequestException类异常,避免程序直接崩溃。 - 设置请求头:模拟浏览器添加
User-Agent,降低被反爬拦截的概率。 - 限制请求频率:添加
time.sleep()延迟,避免短时间内大量请求触发服务器封禁。 - 结构化爬取:爬取时直接按国家维度组织数据为字典结构,减少后续清洗的复杂度。
优化后的爬取代码示例:
import requests from bs4 import BeautifulSoup import json import time def get_country_links(base_url): headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36" } try: response = requests.get(base_url, headers=headers, timeout=10) response.raise_for_status() soup = BeautifulSoup(response.content, 'lxml') country_list = soup.find("div", {"class": "CountryList"}) if not country_list: print("未找到国家列表元素") return [], [] country_links, country_names = [], [] for a_tag in country_list.find_all('a', href=True): country_links.append(a_tag['href']) country_names.append(a_tag.text.strip()) return country_links, country_names except requests.exceptions.RequestException as e: print(f"获取国家列表失败: {e}") return [], [] def scrape_country_data(country_links, country_names, base_url): headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36" } country_data = {} for idx, link in enumerate(country_links[:2]): country_name = country_names[idx] country_data[country_name] = {} full_url = f"{base_url}/{link}" try: response = requests.get(full_url, headers=headers, timeout=10) response.raise_for_status() soup = BeautifulSoup(response.content, 'lxml') menu_list = soup.find('ul', {"class": "pure-menu-list"}) if not menu_list: print(f"{country_name}页面未找到数据列表") continue for li_tag in menu_list.find_all('li'): text = li_tag.text.strip() if not text: continue parts = [p.strip() for p in text.replace('\u00a0', ' ').split('\n') if p.strip()] for i in range(0, len(parts), 2): if i+1 < len(parts): country_data[country_name][parts[i]] = parts[i+1] time.sleep(1) except requests.exceptions.RequestException as e: print(f"爬取{country_name}失败: {e}") continue return country_data if __name__ == "__main__": base_url = "https://data.un.org" country_links, country_names = get_country_links(base_url) if country_links: result = scrape_country_data(country_links, country_names, base_url) with open("structured_country_data.json", 'w', encoding='utf-8') as f: json.dump(result, f, indent=4, ensure_ascii=False) print("结构化数据已保存")
二、数据清洗的更优方案
原始数据的核心特征是“指标名+指标值”成对出现,只是被换行和非断空格干扰。相比简单的字符串拆分,更高效的清洗方式是结构化提取,并做数据标准化处理:
- 替换非断空格,过滤无效空行;
- 按成对规则提取指标与值;
- 可选:去除指标值末尾的标注(如
a/b)、转换数值类型。
针对你现有原始数据的清洗代码示例:
import json # 读取原始数据 with open("data.json", 'r', encoding='utf-8') as f: all_data = json.load(f) cleaned_result = [] for item in all_data: if not item.strip(): continue # 替换非断空格,拆分并过滤空内容 parts = [p.strip() for p in item.replace('\u00a0', ' ').split('\n') if p.strip()] # 跳过标题行 if parts[0] in ["General Information", "Economic indicators", "Social indicators"]: parts = parts[1:] # 成对提取指标和值 country_info = {} for i in range(0, len(parts), 2): if i+1 < len(parts): indicator = parts[i] value = parts[i+1].rstrip('abc') # 去除末尾标注 # 尝试转换数值类型 try: value = float(value.replace(' ', '')) except ValueError: pass country_info[indicator] = value cleaned_result.append(country_info) # 保存清洗后的数据 with open("cleaned_country_data.json", 'w', encoding='utf-8') as f: json.dump(cleaned_result, f, indent=4, ensure_ascii=False) print("数据清洗完成,已保存到cleaned_country_data.json")
内容的提问来源于stack exchange,提问作者basel nabil
相关产品推荐
相关产品推荐

