Python爬取WhatDoTheyKnow网站仅得到表头无数据的问题求助
Python爬取WhatDoTheyKnow网站仅得到表头无数据的问题求助
看起来你的问题主要出在元素选择器不匹配和链接处理不规范这两个核心点上,我帮你一步步分析并修正代码:
问题根源分析
搜索结果页的链接选择器完全错误
你提供的页面源码里,搜索结果的请求链接是嵌套在<span class="head">标签内的<a>元素,而不是你代码里写的a.search-result-heading。这直接导致request_links是空列表,循环逻辑根本没执行,最后自然只有CSV表头,没有数据。相对链接未拼接完整域名
页面里的链接href是相对路径(比如/request/malicious_email_volume_93#incoming-1926496),直接用requests.get(request_url)会请求无效地址,必须拼接网站域名形成完整URL。详情页的元素选择器可能也不匹配
原代码里用h2.name提取问题标题、div.public-description提取回答,这两个选择器和WhatDoTheyKnow详情页的实际结构不符,即使进入详情页也会报错。小问题:重复覆盖
response变量
循环里的response = requests.get(request_url)覆盖了外层请求的response对象,虽然不会直接报错,但会降低代码可读性,建议换个变量名。
修正后的完整代码
# Directory to start in workingdirectory = '/home/webscrape' import requests from bs4 import BeautifulSoup import pandas as pd import time # 配置基础参数 BASE_URL = "https://www.whatdotheyknow.com" url = f"{BASE_URL}/search/O365/all" # 模拟浏览器请求头,避免被反爬拦截 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } # 用Session保持会话,提高请求效率 session = requests.Session() session.headers.update(headers) # Send an HTTP GET request to the URL response = session.get(url) # Check if the request was successful (status code 200) if response.status_code == 200: # Parse the HTML content of the page soup = BeautifulSoup(response.text, "html.parser") # 修正:从span.head下的a标签获取请求链接 request_blocks = soup.find_all("div", class_="request_listing") request_links = [] for block in request_blocks: head_span = block.find("span", class_="head") if head_span: link = head_span.find("a") if link: request_links.append(link) # Initialize lists to store data questions = [] answers = [] # 打印获取到的链接数量,确认是否成功抓取 print(f"共找到 {len(request_links)} 个请求链接") # Loop through each request link and extract data for link in request_links: # 修正:拼接完整URL,去掉锚点部分避免冗余 request_relative_url = link["href"] request_url = BASE_URL + request_relative_url.split("#")[0] try: req_response = session.get(request_url) req_response.raise_for_status() # 主动抛出HTTP请求异常 time.sleep(1) # 添加1秒延迟,避免触发反爬 except requests.exceptions.RequestException as e: print(f"请求详情页失败:{request_url},错误:{e}") continue request_soup = BeautifulSoup(req_response.text, "html.parser") # 修正:提取问题标题(实际详情页标题在h1.request-title) question_elem = request_soup.find("h1", class_="request-title") question = question_elem.text.strip() if question_elem else "未获取到标题" # 修正:提取公开回答内容(找第一个公开可见的回复) answer = "未获取到回答" response_blocks = request_soup.find_all("div", class_="response-block") for block in response_blocks: if "response-visible" in block.get("class", []): content_elem = block.find("div", class_="response-content") if content_elem: answer = content_elem.text.strip() break questions.append(question) answers.append(answer) print(f"已抓取:{question[:50]}...") # Create a DataFrame to store the data data = {"Question": questions, "Answer": answers} df = pd.DataFrame(data) # Save the data to CSV or XLSX df.to_csv("whatdotheyknow_data.csv", index=False) # df.to_excel("whatdotheyknow_data.xlsx", index=False, engine="openpyxl") # 注意原代码的openpyx1是笔误 print("数据抓取并保存成功!") else: print(f"请求搜索页失败,状态码:{response.status_code}") # Close the HTTP session session.close()
额外优化建议
- 分页处理:当前代码仅爬取第一页的25条数据,若需要爬取全部1000条,需实现分页逻辑(比如解析搜索页的下一页链接)。
- 日志记录:可以引入
logging模块替代print,更方便调试和记录问题。 - 随机延迟:把固定的
time.sleep(1)换成随机延迟(比如time.sleep(random.uniform(0.5, 2))),模拟更自然的人类访问行为。
备注:内容来源于stack exchange,提问作者DigiCoinzTrader
相关产品推荐
相关产品推荐

