使用BeautifulSoup在Python中爬取图片URL遇阻求助
问题描述
我编写的爬虫代码在爬取其他网站的图片URL时运行正常,但爬取目标网站https://xx.knit.bid/en/#/?page={shace}时无法获取任何信息,尝试查找该网站的实际请求地址也未成功,代码如下:
import time import requests from bs4 import BeautifulSoup import pandas as pd for i in range(1, 5000): shace=i headers = { "user-agent": "Mozilla/5.0 (Windows NT 5.1) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/44.0.2403.130 Safari/537.36"} url=f'https://xx.knit.bid/en/#/?page={shace}' response = requests.get(url, headers=headers, ) response.encoding = 'utf-8' time.sleep(1) soup = BeautifulSoup(response.text, 'html.parser') for ul in soup.find_all(name='div',attrs={"class":'excerpts'}): div=(ul.find_all(name='article',attrs={"class":'excerpt excerpt-c4'})) hre=[a.get("href")for a in div] title=[a.get("title")for a in div] dit = {'title': title,'URL': hre,} df = pd.DataFrame(dit) df.to_csv(r'D:\88.csv', index=False,mode='a',encoding="utf-8",header=False)
内容的提问来源于stack exchange,提问作者lex9527
相关产品推荐
相关产品推荐

