使用Scrapy爬取h1标签时出现冗余HTML输出的问题
Hey there! The issue you're facing is because your current XPath selector grabs the entire <h1> element (including its tags) instead of just the text inside it. Let's break down how to fix this:
Why You're Seeing Redundant HTML
When you use .get() on resp.xpath('//h1[@class="tdb-title-text"]'), Scrapy returns the full HTML of the matched element. That's why you're getting the entire <h1> tag and its content instead of just the text.
Step 1: Fix the Quotation Mark Issue
First, notice you're using " in your XPath—this is an HTML entity, but in Python strings, you should use regular double quotes (since your string is wrapped in single quotes, double quotes inside work without escaping):
# Instead of " '//h1[@class="tdb-title-text"]'
Step 2: Target the Text Inside the H1
To get just the text content, modify your XPath to target the text node inside the <h1> by adding /text() at the end. Alternatively, you can use Scrapy's CSS selector syntax which is often more readable for text extraction.
Option 1: Using XPath
h1_text = resp.xpath('//h1[@class="tdb-title-text"]/text()').get()
Option 2: Using CSS Selector
h1_text = resp.css('h1.tdb-title-text::text').get()
Both approaches will return only the text inside the <h1> tag, not the surrounding HTML.
Corrected Full Code
Here's your updated code with the fixes applied:
import requests from scrapy.selector import Selector r = requests.get('https://www.catholicgallery.org/mass-reading/310122/') resp = Selector(text=r.text) # Using XPath to get text h1_text = resp.xpath('//h1[@class="tdb-title-text"]/text()').get() # OR using CSS selector # h1_text = resp.css('h1.tdb-title-text::text').get() print(h1_text)
This will print just the text content of the <h1> element without any extra HTML tags.
内容的提问来源于stack exchange,提问作者codewithawais

