Python中如何单独提取div内嵌套元素及指定数据?
可行!用BeautifulSoup轻松提取目标内容
完全可以实现,给你一套适合新手的Python方案,用BeautifulSoup库(爬虫入门常用工具)就能搞定,以下是具体步骤和代码:
步骤说明
- 先安装必要的库(如果还没装):
pip install beautifulsoup4 requests - 解析目标HTML结构,定位到
div#calendar_info_support后,分别提取所需内容:
完整代码示例
from bs4 import BeautifulSoup import re # 你的目标HTML内容(实际爬取时可以用requests.get获取网页内容) html_content = ''' <div id="calendar_info_support"> <a href="http://www.facebook.com/events/" target="_blank">Rescheduled from September 20</a><br> <br> <br> <br> <span class="calendar_info_doors_cover"> 6:30:PM <br> <a href="http://www.showclix.com/event/" target="_blank">$20 advance SC</a> </span> </div> ''' # 解析HTML soup = BeautifulSoup(html_content, 'html.parser') calendar_div = soup.find('div', id='calendar_info_support') # 1. 提取第一个<a>的文本和属性 reschedule_a = calendar_div.find('a') reschedule_text = reschedule_a.get_text(strip=True) reschedule_href = reschedule_a['href'] reschedule_target = reschedule_a['target'] print(f"重排信息文本:{reschedule_text}") print(f"链接地址:{reschedule_href}") print(f"打开方式:{reschedule_target}") # 2. 提取<span>内的时间 time_span = calendar_div.find('span', class_='calendar_info_doors_cover') # 提取span里除子<a>外的文本并清理格式 time_text = ' '.join([text.strip() for text in time_span.stripped_strings if not text.startswith('$')]) time_text = time_text.replace(':PM', ' PM') print(f"活动时间:{time_text}") # 3. 提取第二个<a>里的$20和属性 ticket_a = time_span.find('a') ticket_full_text = ticket_a.get_text(strip=True) ticket_price = re.search(r'\$\d+', ticket_full_text).group() # 精准提取价格 ticket_href = ticket_a['href'] ticket_target = ticket_a['target'] print(f"票价:{ticket_price}") print(f"购票链接:{ticket_href}") print(f"打开方式:{ticket_target}")
代码解释
soup.find():根据标签名、id或class定位元素,新手好上手.get_text(strip=True):提取元素文本并自动去除首尾空格- 提取属性直接用
元素['属性名'],比如a['href']就能拿到链接地址 - 票价用正则
\$\d+匹配,能精准提取$开头的数字金额
内容的提问来源于stack exchange,提问作者Jojeaux
相关产品推荐
相关产品推荐

