You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用requests和lxml爬取网站数据存入CSV仅获第一条数据的解决方法

问题描述

使用requests和lxml模块(仅使用XPath,不可用BeautifulSoup)编写爬虫,目标是从指定页面爬取国家信息并存储到CSV文件,但运行后CSV仅写入第一个国家的信息,需要修正代码实现全量爬取。

原代码如下:

import requests
from lxml import html
import os
import csv
s = requests.session()
headers_dict = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/96.0.4664.45 Safari/537.36"}
#s.headers = headers_dict
r = s.get("https://www.scrapethissite.com/pages/simple/", headers = headers_dict)
tree = html.fromstring(r.content)
rows = tree.xpath('//*[@id="countries"]/div')
print(rows[0].text_content())
# Create a CSV file to store the data
csv_file = open('CountryInfo.csv', 'w')
csv_writer = csv.writer(csv_file)
csv_writer.writerow(['Country', 'Capital', 'Population', 'Area'])
for row in rows:
    country_name = (row.xpath('//*[@id="countries"]/div/div[4]/div[1]/h3'))[0]
    #print(country_name.text_content())
    country_capital = (row.xpath('//*[@id="countries"]/div/div[4]/div[1]/div/span[1]'))[0]
    population = (row.xpath('//*[@id="countries"]/div/div[4]/div[1]/div/span[2]'))[0]
    area = (row.xpath('//*[@id="countries"]/div/div[4]/div[1]/div/span[3]'))[0]
    csv_writer.writerow([country_name.text_content(), country_capital.text_content(), population.text_content(), area.text_content()])
print("Program Executed")
问题根源

循环中使用的XPath是绝对路径(以//开头),每次都会从整个HTML文档的根节点开始查找,导致每次循环都获取到第一个国家的元素,而非当前row节点下的对应子元素。

修正方案

将循环内的XPath改为相对路径(以./开头,代表从当前节点开始查找),或者直接省略//,基于当前row节点定位子元素。同时优化代码细节(比如使用with语句自动管理文件,避免资源泄漏)。

修改后的代码:

import requests
from lxml import html
import csv

s = requests.session()
headers_dict = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/96.0.4664.45 Safari/537.36"}
r = s.get("https://www.scrapethissite.com/pages/simple/", headers=headers_dict)
tree = html.fromstring(r.content)
# 获取所有国家的容器节点
country_containers = tree.xpath('//*[@id="countries"]/div[@class="country"]')

# 使用with语句自动关闭文件
with open('CountryInfo.csv', 'w', newline='', encoding='utf-8') as csv_file:
    csv_writer = csv.writer(csv_file)
    csv_writer.writerow(['Country', 'Capital', 'Population', 'Area'])
    
    for container in country_containers:
        # 相对路径定位当前容器下的元素
        country_name = container.xpath('./h3[@class="country-name"]/text()')[0].strip()
        country_capital = container.xpath('./div[@class="country-info"]/span[@class="country-capital"]/text()')[0].strip()
        population = container.xpath('./div[@class="country-info"]/span[@class="country-population"]/text()')[0].strip()
        area = container.xpath('./div[@class="country-info"]/span[@class="country-area"]/text()')[0].strip()
        
        csv_writer.writerow([country_name, country_capital, population, area])

print("Program Executed")
关键修改说明
  • 循环内XPath改为相对路径:使用./从当前container节点开始查找子元素,确保每次获取的是当前国家的对应信息。
  • 优化元素定位:通过类名(如country-name、country-capital)定位元素,比依赖层级结构更稳定,避免页面布局微调导致XPath失效。
  • 使用with语句管理CSV文件:自动处理文件打开/关闭,防止资源泄漏,同时指定newline=''避免CSV出现多余空行,encoding='utf-8'保证特殊字符正常存储。

内容的提问来源于stack exchange,提问作者Prateek Goyal

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.13 19:50:10