使用requests和lxml爬取网站数据存入CSV仅获第一条数据的解决方法
问题描述
使用requests和lxml模块(仅使用XPath,不可用BeautifulSoup)编写爬虫,目标是从指定页面爬取国家信息并存储到CSV文件,但运行后CSV仅写入第一个国家的信息,需要修正代码实现全量爬取。
原代码如下:
import requests from lxml import html import os import csv s = requests.session() headers_dict = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/96.0.4664.45 Safari/537.36"} #s.headers = headers_dict r = s.get("https://www.scrapethissite.com/pages/simple/", headers = headers_dict) tree = html.fromstring(r.content) rows = tree.xpath('//*[@id="countries"]/div') print(rows[0].text_content()) # Create a CSV file to store the data csv_file = open('CountryInfo.csv', 'w') csv_writer = csv.writer(csv_file) csv_writer.writerow(['Country', 'Capital', 'Population', 'Area']) for row in rows: country_name = (row.xpath('//*[@id="countries"]/div/div[4]/div[1]/h3'))[0] #print(country_name.text_content()) country_capital = (row.xpath('//*[@id="countries"]/div/div[4]/div[1]/div/span[1]'))[0] population = (row.xpath('//*[@id="countries"]/div/div[4]/div[1]/div/span[2]'))[0] area = (row.xpath('//*[@id="countries"]/div/div[4]/div[1]/div/span[3]'))[0] csv_writer.writerow([country_name.text_content(), country_capital.text_content(), population.text_content(), area.text_content()]) print("Program Executed")
问题根源
循环中使用的XPath是绝对路径(以//开头),每次都会从整个HTML文档的根节点开始查找,导致每次循环都获取到第一个国家的元素,而非当前row节点下的对应子元素。
修正方案
将循环内的XPath改为相对路径(以./开头,代表从当前节点开始查找),或者直接省略//,基于当前row节点定位子元素。同时优化代码细节(比如使用with语句自动管理文件,避免资源泄漏)。
修改后的代码:
import requests from lxml import html import csv s = requests.session() headers_dict = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/96.0.4664.45 Safari/537.36"} r = s.get("https://www.scrapethissite.com/pages/simple/", headers=headers_dict) tree = html.fromstring(r.content) # 获取所有国家的容器节点 country_containers = tree.xpath('//*[@id="countries"]/div[@class="country"]') # 使用with语句自动关闭文件 with open('CountryInfo.csv', 'w', newline='', encoding='utf-8') as csv_file: csv_writer = csv.writer(csv_file) csv_writer.writerow(['Country', 'Capital', 'Population', 'Area']) for container in country_containers: # 相对路径定位当前容器下的元素 country_name = container.xpath('./h3[@class="country-name"]/text()')[0].strip() country_capital = container.xpath('./div[@class="country-info"]/span[@class="country-capital"]/text()')[0].strip() population = container.xpath('./div[@class="country-info"]/span[@class="country-population"]/text()')[0].strip() area = container.xpath('./div[@class="country-info"]/span[@class="country-area"]/text()')[0].strip() csv_writer.writerow([country_name, country_capital, population, area]) print("Program Executed")
关键修改说明
- 循环内XPath改为相对路径:使用
./从当前container节点开始查找子元素,确保每次获取的是当前国家的对应信息。 - 优化元素定位:通过类名(如
country-name、country-capital)定位元素,比依赖层级结构更稳定,避免页面布局微调导致XPath失效。 - 使用
with语句管理CSV文件:自动处理文件打开/关闭,防止资源泄漏,同时指定newline=''避免CSV出现多余空行,encoding='utf-8'保证特殊字符正常存储。
内容的提问来源于stack exchange,提问作者Prateek Goyal
相关产品推荐
相关产品推荐

