初学者网页爬虫问题:如何有序抓取NRC反应堆状态页各月份下的每日数据链接
解决方案
你可以通过定位h2标签后遍历后续兄弟节点的方式实现结构化抓取,逻辑为:每匹配到一个月份h2标签后,持续收集其后续的兄弟元素,直到遇到下一个h2标签停止,中间的所有有效日度链接都归属当前月份,最终生成嵌套的结构化数据结构,方便存储。
修正后的可运行代码如下:
from bs4 import BeautifulSoup import requests # 替换为你需要抓取的目标年份页面 base_url = "https://www.nrc.gov/reading-rm/doc-collections/event-status/reactor-status/2004/index.html" domain_prefix = "https://www.nrc.gov" resp = requests.get(base_url) soup = BeautifulSoup(resp.text, "html.parser") # 初始化结构化存储字典:key为月份,value为对应月份的所有日度数据链接 monthly_links = {} current_month = None # 遍历页面所有h2标签(对应月份锚点) for h2 in soup.find_all('h2'): current_month = h2.get_text(strip=True) monthly_links[current_month] = [] # 遍历当前h2的所有后续兄弟节点 next_sibling = h2.find_next_sibling() while next_sibling: # 遇到下一个h2标签时终止当前月份的链接收集 if next_sibling.name == 'h2': break # 只收集有效a标签链接,排除锚点、当前页等无效链接 if next_sibling.name == 'a' and next_sibling.get('href'): href = next_sibling['href'] # 补全相对路径为绝对路径 if not href.startswith('http'): href = domain_prefix + href # 排除页面顶部的月份锚点链接 if '#' not in href: monthly_links[current_month].append(href) next_sibling = next_sibling.find_next_sibling() # 打印结构化结果,也可以直接转JSON/CSV存储 for month, links in monthly_links.items(): print(f"{month} 共 {len(links)} 条日度数据链接:") for link in links: print(f"- {link}")
结果说明
最终得到的monthly_links是嵌套结构,直接对应「月份-日度链接」的层级关系,你可以根据需求直接转成JSON文件存储,或者批量下载对应链接的文件即可。
内容的提问来源于stack exchange,提问作者RCoder
相关产品推荐
相关产品推荐

