Python网页爬虫无警告无报错无输出问题,使用requests与BeautifulSoup
代码无输出无报错的问题原因及解决方案
问题根因
soup.find_all()调用逻辑错误:find_all方法的第一个参数是HTML标签名,你传入col-md-2会被识别为查找<col-md-2>标签,而你的目标是查找class为col-md-2的div元素,匹配不到任何内容时后续代码完全不执行,因此没有输出也没有报错。- 文本获取属性错误:代码中
pictags.txt是无效属性,BeautifulSoup标签对象获取文本的正确属性是.text,也可以调用.get_text()方法。 - 缺少反爬兼容与状态校验:默认
requests.get()请求的UA标识为python-requests,绝大多数网站会直接拦截该类爬虫请求,返回异常状态码,你没有对请求结果做校验,即使请求失败也不会有提示。 - 缺少文件内容校验:如果
links.txt为空、或存储的链接不是合法HTTP链接,也会导致无执行结果。
修复方案
- 调整
find_all的写法,指定筛选class为col-md-2的div元素 - 替换文本获取的属性为
.text - 新增自定义请求头,模拟浏览器请求,新增状态码校验逻辑
- 新增
links.txt非空判断
修复后完整代码
import requests import random import string from bs4 import BeautifulSoup # 模拟浏览器的请求头,避免被反爬拦截 HEADERS = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } # Create random string of specific length def randStr(chars = string.ascii_uppercase + string.digits, N=10): return ''.join(random.choice(chars) for _ in range(N)) if __name__ == "__main__": with open("links.txt", "r", encoding="utf-8") as a_file: lines = [line.strip() for line in a_file if line.strip()] if not lines: print("links.txt无有效链接") exit() for endpoint in lines: # 跳过非http开头的无效链接 if not endpoint.startswith("http"): continue try: response = requests.get(endpoint, headers=HEADERS, timeout=10) # 校验请求是否成功 response.raise_for_status() except Exception as e: print(f"请求{endpoint}失败:{e}") continue soup = BeautifulSoup(response.text, "html.parser") # 调整find_all写法,筛选class为col-md-2的div元素 for pictags in soup.find_all("div", class_="col-md-2"): lastfilename = randStr() # 新增编码指定,避免中文乱码 with open(f"{lastfilename}.txt", "w", encoding="utf-8") as file: # 用正确的text属性获取文本 file.write(pictags.text.strip()) print(f"已保存{endpoint}的内容到{lastfilename}.txt")
内容的提问来源于stack exchange,提问作者Katherine Elizabeth Kath
相关产品推荐
相关产品推荐

