Python网页爬取遇相同页面HTML结构差异问题求助
解决网页爬取中HTML结构不一致导致程序崩溃的问题
核心思路
先判断当前页面的结构类型,针对不同结构执行对应逻辑,同时给表格提取逻辑加上容错处理,避免程序直接终止。
具体实现步骤
结构检测:
通过检查页面中是否存在目标表格的标志性元素(比如表格的id、class,或特定表头文本),或是异常结构的特征(比如包含错误提示的脚本内容)来区分结构。示例代码:import requests from bs4 import BeautifulSoup url = "https://www.ibm.com/docs/en/imdm/12.0?topic=t-accessdateval" # 模拟浏览器请求头,降低异常结构出现概率 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36", "Accept-Language": "zh-CN,zh;q=0.9" } response = requests.get(url, headers=headers) soup = BeautifulSoup(response.text, "html.parser") # 检测异常结构:替换成实际异常脚本中的关键词 is_exception = False for script in soup.find_all("script"): if script.string and "错误提示翻译" in script.string: is_exception = True break # 检测目标表格:替换成实际表格的class/id target_table = soup.find("table", class_="ibm-table")分支处理逻辑:
- 若为异常结构:直接跳过或记录日志后继续执行后续任务;
- 若为预期结构:执行表格提取和CSV保存,同时用
try-except包裹核心代码,处理提取时的意外错误。
示例代码:
import csv if is_exception: print(f"页面 {url} 为异常结构,跳过") with open("crawl_error.log", "a", encoding="utf-8") as f: f.write(f"{url} - 异常结构\n") elif target_table: try: # 提取表头 headers = [th.get_text(strip=True) for th in target_table.find_all("th")] # 提取表格内容行 rows = [] for tr in target_table.find_all("tr")[1:]: row_data = [td.get_text(strip=True) for td in tr.find_all("td")] rows.append(row_data) # 追加写入CSV with open("table_data.csv", "a", newline="", encoding="utf-8") as csvfile: writer = csv.writer(csvfile) # 仅当文件为空时写入表头 if csvfile.tell() == 0: writer.writerow(headers) writer.writerows(rows) print(f"成功处理页面 {url}") except Exception as e: print(f"处理页面 {url} 出错: {str(e)}") with open("crawl_error.log", "a", encoding="utf-8") as f: f.write(f"{url} - 提取错误: {str(e)}\n") else: print(f"页面 {url} 未找到目标表格,跳过") with open("crawl_error.log", "a", encoding="utf-8") as f: f.write(f"{url} - 无目标表格\n")额外优化:
- 加入请求重试机制:使用
requests.adapters.HTTPAdapter设置重试次数,应对临时请求异常; - 定期清理日志文件,避免日志过大。
- 加入请求重试机制:使用
内容的提问来源于stack exchange,提问作者Agreeable_Math_314
相关产品推荐
相关产品推荐

