Python BeautifulSoup爬取黄页电话号码返回None问题求助
问题原因
电话字段返回None由3个代码错误导致:
- 电话赋值行残留了无效占位文本
enter code here,会直接导致匹配逻辑失效甚至语法报错 - 仅匹配到了外层class为
popover-phones的div节点,没有继续向下查找存储实际号码的a标签 - 没有提取a标签的
href属性,也未对号码前缀做清洗,即便找到节点也拿不到可用的号码文本
另外你代码开头的import item as item、import json属于未使用的冗余导入,可直接删除。
修正逻辑
按页面实际结构调整提取规则,同时兼容单号码、多号码场景,避免无号码条目触发报错:
- 先在条目节点下查找
popover-phonesdiv,做非空判断 - 找到div后提取内部所有a标签
- 遍历a标签取
href属性,切掉开头的tel:前缀得到纯号码 - 多个号码用逗号拼接,方便后续存储
修正后可运行代码
import requests from bs4 import BeautifulSoup from csv import writer url = 'https://yellowpages.com.eg/en/category/charcoal' headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/103.0.0.0 Safari/537.36'} r = requests.get(url, headers=headers) soup = BeautifulSoup(r.content, 'html.parser') articles = soup.find_all('div', class_='col-xs-12 item-details') for item in articles: address = item.find('a', class_='address-text').text.strip() company = item.find('a', class_='item-title').text.strip() telephone = '' phone_div = item.find('div', class_='popover-phones') if phone_div: phone_list = [] for a_tag in phone_div.find_all('a'): tel_link = a_tag.get('href', '') if tel_link.startswith('tel:'): phone_list.append(tel_link.replace('tel:', '').strip()) telephone = ','.join(phone_list) print(company, address, telephone)
补充:如果修正后还是匹配不到
popover-phones节点,说明该节点是JS动态渲染生成,普通requests请求拿到的静态源码里不存在该节点,这种情况要么直接抓站点加载号码的后端接口,要么改用Selenium、Pyppeteer这类可执行JS的渲染工具抓取即可。
内容的提问来源于stack exchange,提问作者Fouad Elmahdy
相关产品推荐
相关产品推荐

