如何使用Beautiful Soup提取网页p标签数据 运行无返回结果如何解决
代码问题说明
- 首次请求阿拉巴马州站点列表页
https://locations.atipt.com/al时未携带请求头,触发站点反爬规则,返回非有效页面内容,无法匹配到list-unstyled类的ul标签,导致productlinks列表为空,后续遍历逻辑完全不执行,无输出。 - 详情页标签提取逻辑错误:目标页面
https://locations.atipt.com/al/alabaster中,name类的a标签下不存在子级p标签,门店地址、联系方式等p标签和该a标签为同级节点,原有提取路径无法匹配到目标元素。
调整后可运行代码
import requests from bs4 import BeautifulSoup import pandas as pd baseurl = 'https://locations.atipt.com/' headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.114 Safari/537.36' } # 首次请求也携带headers避免被反爬拦截 r = requests.get('https://locations.atipt.com/al', headers=headers) soup = BeautifulSoup(r.content, 'html.parser') tra = soup.find_all('ul', class_='list-unstyled') productlinks = [] for links in tra: for link in links.find_all('a', href=True): comp = baseurl + link['href'] productlinks.append(comp) # 遍历详情页提取内容 for link in productlinks: r = requests.get(link, headers=headers) soup = BeautifulSoup(r.content, 'html.parser') # 直接从content-card容器提取所有p标签 tag = soup.find_all('div', class_='listing content-card') for pro in tag: p_list = pro.find_all('p') for p in p_list: print(p.get_text(strip=True)) # 不同门店内容加分隔线方便区分 print("="*30)
补充说明
如果仅需要提取单个目标页https://locations.atipt.com/al/alabaster的内容,可以跳过列表页爬取逻辑,直接请求该详情页执行p标签提取即可。
内容的提问来源于stack exchange,提问作者Amen Aziz
相关产品推荐
相关产品推荐

