爬取网页后如何匹配关联标题与对应时间并按要求结构化存储
代码调整方案
核心思路
不要全局分开抓取所有标题和时间元素,优先以每个会议的外层共同容器为单位提取内容,从根源上避免两个字段顺序错位的问题,保证时间和标题一一对应。
完整可运行代码
from bs4 import BeautifulSoup from selenium import webdriver import time driver = webdriver.Chrome() driver.get('https://ash.confex.com/ash/2021/webprogram/STUDIO.html') time.sleep(3) page_source = driver.page_source soup = BeautifulSoup(page_source, 'html.parser') driver.quit() # 定位所有会议条目的外层容器 meeting_blocks = soup.find_all('div', class_='item') output_list = [] for block in meeting_blocks: # 提取当前条目的时间 time_tag = block.find('span', class_='time') if not time_tag: continue meet_time = time_tag.text.strip() # 提取当前条目的标题 title_tag = block.find('div', class_='itemtitle') if not title_tag: continue a_tag = title_tag.find('a', href=True) if not a_tag: continue meet_title = a_tag.text.strip() # 按要求格式拼接 output_list.append(f"{meet_time}, {meet_title}") # 逐行输出结果 for line in output_list: print(line) # 如需存入本地文件可取消下方注释 # with open('ash_meeting.csv', 'w', encoding='utf-8') as f: # f.write('\n'.join(output_list))
轻量替代方案(仅适合元素顺序严格对应的场景)
如果确认页面内所有时间元素和标题元素的排列顺序完全匹配,也可以用zip方法直接配对两个列表,写法更简洁:
from bs4 import BeautifulSoup from selenium import webdriver import time driver = webdriver.Chrome() driver.get('https://ash.confex.com/ash/2021/webprogram/STUDIO.html') time.sleep(3) page_source = driver.page_source soup = BeautifulSoup(page_source, 'html.parser') driver.quit() # 收集所有标题 titles = [] for title_block in soup.find_all('div', class_='itemtitle'): a_tag = title_block.find('a', href=True) if a_tag: titles.append(a_tag.text.strip()) # 收集所有时间 times = [] for time_tag in soup.find_all('span', class_='time'): times.append(time_tag.text.strip()) # 配对输出 for t, title in zip(times, titles): print(f"{t}, {title}")
内容的提问来源于stack exchange,提问作者Age
相关产品推荐
相关产品推荐

