You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python爬虫使用Chrome XPath无法获取网站列表标题问题求助

解决爬虫主体容器返回空列表的问题

Hey there! I’ve dealt with this exact frustration way too many times—when some XPaths work (like the search box text) but your main container returns [] every time. Let’s break down the most likely fixes, starting with the most common culprits:

1. 动态内容渲染(90%的概率是这个)

Most modern sites load content dynamically with JavaScript, but the requests library only fetches the static initial HTML sent by the server. The content you see in your browser’s dev tools is built after the page loads via JS, so requests can’t see it.

Fix: Use a tool that simulates a real browser

Tools like Selenium or Playwright will render the page just like a human’s browser, so you’ll get the fully loaded content. Here’s a quick Selenium example to replace your requests call:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

# Initialize browser (make sure you have ChromeDriver installed)
driver = webdriver.Chrome()
driver.get("https://your-target-site.com")

# Wait for the main container to load (avoids race conditions)
wait = WebDriverWait(driver, 10)
main_container = wait.until(EC.presence_of_element_located((By.XPATH, "your-main-container-xpath")))

# Grab titles using a relative XPath (note the dot at the start!)
titles = main_container.find_elements(By.XPATH, ".//your-title-xpath")

for title in titles:
    print(title.text)

driver.quit()

2. 服务器给requests返回了不同的HTML

Sometimes sites serve different content to bots vs. real browsers. This could be due to:

  • Missing User-Agent header: Add a real browser’s UA to your requests call:
    headers = {
        "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
    }
    response = requests.get("your-url", headers=headers)
    
  • Missing cookies: Some sites require you to first visit the homepage to get session cookies, then reuse them for the list page.

Quick check: Print response.text and search for your target title. If it’s not there, the server isn’t sending it to requests—so dynamic content or anti-bot measures are to blame.

3. 目标内容嵌套在iframe里

If the main container lives inside an <iframe> tag, requests won’t load its content automatically. You’ll need to:

  1. Parse the initial HTML to find the iframe’s src attribute.
  2. Make a separate requests call (or use Selenium to switch to the iframe) to fetch that content.

4. 反爬机制拦截了请求

Some sites block repeated requests or detect bots. Try:

  • Adding delays between requests with time.sleep(2).
  • Rotating user agents or using a proxy if you’re getting blocked.

Pro Tip: 检查浏览器Network标签

Open your browser’s dev tools → Network tab → 刷新页面。寻找加载列表数据的XHR/Fetch请求。通常你可以直接调用这些返回JSON的API,比解析HTML更快更可靠!

内容的提问来源于stack exchange,提问作者Ke Wang

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 04:13:11