You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Scrapy抓取页面全部文本?response.css()如何选择所有标签?

关于Scrapy抓取所有文本的解决方案

嘿,刚好能帮到你!在Scrapy的response.css()里,确实有办法选中所有标签的文本——用CSS通配符*搭配::text伪元素就行,直接就能提取页面上所有标签里的文本内容。

核心用法说明

要提取页面所有文本,最直接的CSS选择器写法是:

all_text_nodes = response.css('*::text').getall()

不过这个写法会拿到很多空白字符(比如换行、连续空格),所以一定要做下清洗,过滤掉无效内容:

cleaned_text = [text.strip() for text in all_text_nodes if text.strip()]
full_text = ' '.join(cleaned_text)

这样就能得到干净的纯文本内容了。

结合你的代码修改示例

看你提到要跟进部分链接,用CrawlSpider配合链接规则会比普通Spider更省心,我基于你的代码调整了一下,直接能用:

import scrapy
from bs4 import BeautifulSoup
import nltk
import lxml.html
import pandas as pd
from scrapy.spiders import CrawlSpider, Rule
from scrapy.linkextractors import LinkExtractor

class QuotesSpider(CrawlSpider):
    name = "dialpad"
    # 一定要设置允许爬取的域名,避免爬取无关网站
    allowed_domains = ["你的目标域名.com"]
    # 替换成你的起始页面URL
    start_urls = ["https://你的起始页面.com"]

    # 定义链接跟进规则:只爬取目标域名内的链接,排除静态资源(css/js/图片等)
    rules = (
        Rule(
            LinkExtractor(
                allow=r'^https://你的目标域名\.com/',
                deny=r'\.(css|js|png|jpg|gif)$'
            ),
            callback='parse_page',
            follow=True
        ),
    )

    def parse_page(self, response):
        # 提取所有文本并清洗
        all_text = response.css('*::text').getall()
        cleaned_text = [t.strip() for t in all_text if t.strip()]
        full_text = ' '.join(cleaned_text)

        # 输出结果,可以按需保存到文件、数据库或Scrapy Item
        yield {
            'url': response.url,
            'full_content': full_text
        }

补充小提示

  • 如果你习惯用XPath,也可以用response.xpath('//text()').getall(),效果和CSS写法完全一致,选你顺手的就行
  • 记得根据实际需求调整LinkExtractor的allow和deny规则,精准控制要跟进的链接
  • 如果不想用CrawlSpider,用普通Spider的话,就在parse方法里手动遍历链接并调用response.follow()跟进就行

内容的提问来源于stack exchange,提问作者Pranav Barot

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 04:15:46