You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用BeautifulSoup从指定网页抓取姓名并存储到数组?

解决BeautifulSoup抓取指定页面男孩姓名的问题

适配两个页面的选择器

针对你提供的两个页面,需使用对应CSS选择器定位姓名元素:

  1. 双胞胎男孩名页面:.name-pair .name —— 每个名字对里的单个姓名均匹配该选择器
  2. 现代印度教男孩名页面:.name-list-item .name —— 所有姓名均包含在该选择器匹配的标签中

修改后的完整代码

from bs4 import BeautifulSoup
import requests

# 目标URL列表
urls = [
    "https://angelsname.com/Twin-Boy-Names",
    "https://angelsname.com/Modern-Hindu-Baby-Names/Boy"
]

names = []

for url in urls:
    # 模拟浏览器请求,避免被拦截
    headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"}
    html = requests.get(url, headers=headers)
    soup = BeautifulSoup(html.text, 'lxml')
    
    # 根据URL匹配对应选择器
    selectors = ".name-pair .name" if "Twin-Boy-Names" in url else ".name-list-item .name"
    
    # 提取姓名文本并去重
    for result in soup.select(selectors):
        name_text = result.get_text(strip=True)
        if name_text and name_text not in names:
            names.append(name_text)

# 输出结果
print(names)

关键修正说明

  • 修复原代码语法错误:补充for循环末尾冒号,给select()参数添加引号
  • 添加请求头:避免网站将请求识别为爬虫返回空内容
  • 提取纯文本:用get_text(strip=True)获取干净的姓名内容,而非标签对象
  • 可选去重:避免双胞胎页面出现重复姓名,可根据需求删除判断条件

内容的提问来源于stack exchange,提问作者NooB Programmer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 22:30:55