Python爬取OLX页面解析广告标题报AttributeError问题求解
OLX爬取广告标题报错排查方案
报错原因
你遇到的AttributeError: 'NoneType' object has no attribute 'get_text'由三个问题共同导致:
- 语法参数写错:BeautifulSoup按class查找元素的参数名是
class_(单下划线),你的代码里写的是class__(双下划线),导致标题节点匹配逻辑完全失效,find()方法返回空值None,后续调用get_text()直接触发报错。 - 缺少容错判断:就算修正参数名,
css-19ucd76匹配到的div块里会混杂推广占位、页面组件等非广告内容,这些块内部不存在目标标题节点,直接取文本同样会触发空对象报错。 - 缩进语法错误:原代码中
return animal的缩进层级错误,不在get_content函数的内部作用域,运行时还会触发缩进相关的语法报错。
修正后可运行代码
import requests from bs4 import BeautifulSoup import csv HOST = 'https://www.olx.ua/' URL = 'https://www.olx.ua/d/zhivotnye/sobaki/' HEADERS = { 'accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,image/apng,*/*;q=0.8,application/signed-exchange;v=b3;q=0.9', 'user-agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/102.0.0.0 Safari/537.36' } def get_html(url, params=''): r = requests.get(url, headers=HEADERS, params=params) return r def get_content(html): soup = BeautifulSoup(html, 'html.parser') items = soup.find_all('div', class_='css-19ucd76') animal = [] for item in items: title_node = item.find('div', class_='css-u2ayx9') # 跳过无标题的无效块 if not title_node: continue animal.append( { 'Title': title_node.get_text(strip=True) } ) return animal if __name__ == "__main__": html = get_html(URL) if html.status_code == 200: ad_list = get_content(html.text) print(ad_list) # 后续写入CSV的逻辑可直接在此处补充 else: print(f"页面请求失败,状态码:{html.status_code}")
后续爬取注意事项
- OLX的前端class类名是随版本迭代动态生成的,如果后续运行出现匹配不到内容的情况,优先打开浏览器开发者工具核对当前页面实际的class属性值,不要硬编码旧的类名。
- 写入CSV时直接使用标准库的
csv.DictWriter,传入爬取得到的广告列表即可批量写入,不需要额外做格式转换。
内容的提问来源于stack exchange,提问作者Akatsyki
相关产品推荐
相关产品推荐

