Python爬取指定class的div内p标签内容无输出如何解决?
问题背景
需要通过Python爬取日语每日一词网站的当日词汇及对应释义,自行编写的爬虫代码运行后无任何输出,无法拿到目标内容。
网页参考截图:
原有问题代码如下:
import requests from bs4 import BeautifulSoup html_text = requests.get('https://www.transparent.com/word-of-the-day/today/japanese.html').text soup = BeautifulSoup(html_text,'html.parser') content =soup.find_all('div',class_='wotdr-item wotdr-item--translation js-col-item') for i in content: for p in i.find_all('p'): print(p.text)
问题原因
- 未添加浏览器请求头:requests默认的请求UA会被站点的基础反爬机制识别,直接返回非完整渲染的异常页面,目标DOM节点根本不存在于返回的源码中。
- 选择器使用了动态类名:代码中匹配的
js-col-item是前端JS逻辑绑定用的动态类,部分场景下不会在初始返回的HTML中存在,且全量匹配多类名的写法容错性极低,只要站点前端微调类名就会匹配失效。
修正后可运行代码
import requests from bs4 import BeautifulSoup # 模拟正常浏览器的请求头,绕过基础反爬校验 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/125.0.0.0 Safari/537.36" } response = requests.get( 'https://www.transparent.com/word-of-the-day/today/japanese.html', headers=headers ) response.encoding = "utf-8" soup = BeautifulSoup(response.text, "html.parser") # 提取当日核心词汇 head_word = soup.find("div", class_="wotd-head-word").text.strip() print(f"今日单词:{head_word}") # 提取释义内容,使用稳定的静态样式类匹配,跳过动态JS绑定类 translation_content = soup.find("div", class_="wotdr-item--translation") if translation_content: for p_tag in translation_content.find_all("p"): text = p_tag.text.strip() if text: print(text)
额外说明
- 若运行后仍匹配不到内容,可以先打印
response.text查看返回的源码内容,确认是否被反爬拦截,必要时可补充Referer、Cookie等请求头字段。 - 编写爬虫选择器时,尽量避开带
js-前缀的类名,这类类名服务于前端交互逻辑,变动概率远高于静态样式类,会大幅提升代码维护成本。
内容的提问来源于stack exchange,提问作者ikichiziki
相关产品推荐
相关产品推荐

