如何用Selenium(Python)从XPath提取多元素并写入Excel单行?
Hey there! Let's break down how to solve your three problems with Selenium and Excel handling, step by step:
Your original code uses find_element() which only returns the first matching element. To grab all elements that fit your XPath, switch to find_elements() (note the plural), then extract text from each element in the resulting list.
Here's the adjusted code:
from selenium import webdriver from selenium.webdriver.common.by import By driver = webdriver.Chrome() driver.get('http://www.imdb.com/title/tt4854442/?ref_=wl_li_tt') # 提取所有匹配元素的文本,存入列表 producers = [elem.text for elem in driver.find_elements(By.XPATH, "//div[@class='txt-block']/span/a/span")] print(producers) # 输出会是包含所有制片人名称的列表
Your original XPath is too generic—txt-block is used for multiple sections on IMDb pages (like directors, writers, and more). To target only producers, we can anchor the path to the "Producers" section header for better accuracy:
//h4[text()='Producers']/following-sibling::div/span/a/span
Why this works:
//h4[text()='Producers']locates the exact header for the producers section/following-sibling::divselects the sibling div right after the header (where the producer names are stored)- The rest of the path targets the name links inside that specific div
Updated code with the optimized XPath:
producers = [elem.text for elem in driver.find_elements(By.XPATH, "//h4[text()='Producers']/following-sibling::div/span/a/span")]
There are two common approaches here—either put all producers in a single cell (separated by commas) or place each producer in its own column on the same row.
Option 1: All producers in one cell
import xlwt # 创建工作簿和工作表 wb = xlwt.Workbook() ws = wb.add_sheet("Producers") # 将列表转为逗号分隔的字符串 producers_str = ', '.join(producers) # 写入第0行第0列(Excel行/列从0开始计数) ws.write(0, 0, producers_str) # 保存文件 wb.save("imdb_producers.xls")
Option 2: Each producer in a separate column on the same row
import xlwt wb = xlwt.Workbook() ws = wb.add_sheet("Producers") # 遍历列表,逐个写入同一行的不同列 for col_idx, producer in enumerate(producers): ws.write(0, col_idx, producer) wb.save("imdb_producers.xls")
Full combined code (end-to-end workflow)
Here's the complete flow from scraping to saving:
from selenium import webdriver from selenium.webdriver.common.by import By import xlwt driver = webdriver.Chrome() driver.get('http://www.imdb.com/title/tt4854442/?ref_=wl_li_tt') # 用优化后的XPath获取所有制片人 producers = [elem.text for elem in driver.find_elements(By.XPATH, "//h4[text()='Producers']/following-sibling::div/span/a/span")] # 写入Excel(这里用多列示例,可替换为单单元格写法) wb = xlwt.Workbook() ws = wb.add_sheet("Producers") for col_idx, producer in enumerate(producers): ws.write(0, col_idx, producer) wb.save("imdb_producers.xls") driver.quit()
内容的提问来源于stack exchange,提问作者Luís Henrique Martins

