You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python爬取指定class的div内p标签内容无输出如何解决?

问题背景

需要通过Python爬取日语每日一词网站的当日词汇及对应释义,自行编写的爬虫代码运行后无任何输出,无法拿到目标内容。
网页参考截图:
网页目标内容参考截图

原有问题代码如下:

import requests
from bs4 import BeautifulSoup

html_text = requests.get('https://www.transparent.com/word-of-the-day/today/japanese.html').text
soup = BeautifulSoup(html_text,'html.parser')
content =soup.find_all('div',class_='wotdr-item wotdr-item--translation js-col-item')
for i in content:
   for p in i.find_all('p'):
       print(p.text)
问题原因
  • 未添加浏览器请求头:requests默认的请求UA会被站点的基础反爬机制识别,直接返回非完整渲染的异常页面,目标DOM节点根本不存在于返回的源码中。
  • 选择器使用了动态类名:代码中匹配的js-col-item是前端JS逻辑绑定用的动态类,部分场景下不会在初始返回的HTML中存在,且全量匹配多类名的写法容错性极低,只要站点前端微调类名就会匹配失效。
修正后可运行代码
import requests
from bs4 import BeautifulSoup

# 模拟正常浏览器的请求头,绕过基础反爬校验
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/125.0.0.0 Safari/537.36"
}

response = requests.get(
    'https://www.transparent.com/word-of-the-day/today/japanese.html',
    headers=headers
)
response.encoding = "utf-8"
soup = BeautifulSoup(response.text, "html.parser")

# 提取当日核心词汇
head_word = soup.find("div", class_="wotd-head-word").text.strip()
print(f"今日单词:{head_word}")

# 提取释义内容,使用稳定的静态样式类匹配,跳过动态JS绑定类
translation_content = soup.find("div", class_="wotdr-item--translation")
if translation_content:
    for p_tag in translation_content.find_all("p"):
        text = p_tag.text.strip()
        if text:
            print(text)
额外说明
  • 若运行后仍匹配不到内容,可以先打印response.text查看返回的源码内容,确认是否被反爬拦截,必要时可补充Referer、Cookie等请求头字段。
  • 编写爬虫选择器时,尽量避开带js-前缀的类名,这类类名服务于前端交互逻辑,变动概率远高于静态样式类,会大幅提升代码维护成本。

内容的提问来源于stack exchange,提问作者ikichiziki

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.28 22:54:11