如何使用BeautifulSoup从指定网页抓取姓名并存储到数组?
解决BeautifulSoup抓取指定页面男孩姓名的问题
适配两个页面的选择器
针对你提供的两个页面,需使用对应CSS选择器定位姓名元素:
- 双胞胎男孩名页面:
.name-pair .name—— 每个名字对里的单个姓名均匹配该选择器 - 现代印度教男孩名页面:
.name-list-item .name—— 所有姓名均包含在该选择器匹配的标签中
修改后的完整代码
from bs4 import BeautifulSoup import requests # 目标URL列表 urls = [ "https://angelsname.com/Twin-Boy-Names", "https://angelsname.com/Modern-Hindu-Baby-Names/Boy" ] names = [] for url in urls: # 模拟浏览器请求,避免被拦截 headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"} html = requests.get(url, headers=headers) soup = BeautifulSoup(html.text, 'lxml') # 根据URL匹配对应选择器 selectors = ".name-pair .name" if "Twin-Boy-Names" in url else ".name-list-item .name" # 提取姓名文本并去重 for result in soup.select(selectors): name_text = result.get_text(strip=True) if name_text and name_text not in names: names.append(name_text) # 输出结果 print(names)
关键修正说明
- 修复原代码语法错误:补充
for循环末尾冒号,给select()参数添加引号 - 添加请求头:避免网站将请求识别为爬虫返回空内容
- 提取纯文本:用
get_text(strip=True)获取干净的姓名内容,而非标签对象 - 可选去重:避免双胞胎页面出现重复姓名,可根据需求删除判断条件
内容的提问来源于stack exchange,提问作者NooB Programmer
相关产品推荐
相关产品推荐

