Python网页爬取遇空列表及TypeError问题求助
解决Python网页爬取的两个问题
问题1:代码返回空数据
原因分析
- 无用的
from urllib import response导入可能导致变量冲突,且完全不需要。 - 网站可能拦截了无浏览器标识的请求,返回空内容。
- 目标内容可能是JavaScript动态渲染的,
requests只能获取静态HTML。
解决方案
方案1:添加请求头模拟浏览器
修改代码,移除无用导入并添加请求头:
import requests from bs4 import BeautifulSoup import pandas as pd kitapurl = "https://1000kitap.com/alintilar" # 模拟Chrome浏览器请求头 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36" } response = requests.get(kitapurl, headers=headers) # 先验证请求是否成功(200表示正常) print(response.status_code) soup = BeautifulSoup(response.content, "html.parser") gelen_ana_veri = soup.find_all('span', attrs={'class':'text-alt'}) print(gelen_ana_veri)
方案2:处理动态渲染内容(如果方案1无效)
若内容是JS动态加载的,使用Selenium获取渲染后的页面:
from selenium import webdriver from selenium.webdriver.chrome.service import Service from bs4 import BeautifulSoup import pandas as pd kitapurl = "https://1000kitap.com/alintilar" # 替换为你的chromedriver路径 service = Service("chromedriver.exe") driver = webdriver.Chrome(service=service) driver.get(kitapurl) # 获取渲染后的页面源码 soup = BeautifulSoup(driver.page_source, "html.parser") gelen_ana_veri = soup.find_all('span', attrs={'class':'text-alt'}) print(gelen_ana_veri) driver.quit()
问题2:触发TypeError错误
原因分析
soup.find_all()返回的是元素列表,你直接用字符串'content'索引列表,违反了列表只能用整数/切片索引的规则,因此报错。
解决方案
方法1:使用find()获取单个元素(推荐)
find()返回匹配到的第一个元素,直接获取其属性:
import requests from bs4 import BeautifulSoup import pandas as pd kitapurl = "https://1000kitap.com/alintilar" headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36" } response = requests.get(kitapurl, headers=headers) soup = BeautifulSoup(response.content, "html.parser") # 安全获取属性:用get()避免元素不存在时抛出异常 meta_tag = soup.find('meta', attrs={'property':'og:description'}) if meta_tag: gelen_ana_veri = meta_tag.get('content') print(gelen_ana_veri) else: print("未找到目标meta标签")
方法2:从find_all()结果中取第一个元素
如果必须用find_all(),先判断列表非空再索引:
import requests from bs4 import BeautifulSoup import pandas as pd kitapurl = "https://1000kitap.com/alintilar" headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36" } response = requests.get(kitapurl, headers=headers) soup = BeautifulSoup(response.content, "html.parser") meta_tags = soup.find_all('meta', attrs={'property':'og:description'}) if meta_tags: gelen_ana_veri = meta_tags[0].get('content') print(gelen_ana_veri) else: print("未找到目标meta标签")
内容的提问来源于stack exchange,提问作者Fırat Kaya
相关产品推荐
相关产品推荐

