You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python新手爬取Wikifolio网站portfolio表格失败求助

Troubleshooting Your Wikifolio Portfolio Scraping Issue

Hey Julian, I totally get the frustration—spending hours testing 15+ methods and still hitting a wall with that portfolio data must be so annoying. Let’s break down why your current code isn’t working and fix it step by step.

Why Your Current Methods Fail

The core issue here is dynamic content loading. When you use requests.get() or urlopen() to fetch the page, you’re only getting the initial HTML sent by the server. The "Portfolio" table you want doesn’t exist in that initial HTML—it’s loaded after the page loads, via JavaScript pulling data from an internal API and rendering it on the page. That’s why your code keeps grabbing the static right-side table instead: it’s present from the start, while the portfolio rows aren’t.

Fix 1: Scrape the Direct API (Most Efficient)

Instead of wrestling with HTML parsing, you can pull the portfolio data directly from the API the website uses. This is faster and gives you clean, structured data.

Step 1: Find the API Endpoint

  1. Open the Wikifolio page in Chrome/Firefox, hit F12 to open DevTools.
  2. Go to the Network tab, refresh the page.
  3. Filter by "XHR" to narrow down requests—you’ll see a call that returns the portfolio data. For your page, it’s likely something like https://www.wikifolio.com/api/de/de/instrument/wffalkinve/portfolio.

Step 2: Code to Fetch from the API

import requests
import pandas as pd

# Replace with the actual API endpoint you found in DevTools
API_URL = "https://www.wikifolio.com/api/de/de/instrument/wffalkinve/portfolio"
HEADERS = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
}

# Fetch the data
response = requests.get(API_URL, headers=HEADERS)
response.raise_for_status()  # Raise error if request fails
data = response.json()

# Extract portfolio items and convert to DataFrame
portfolio_items = data["data"]["items"]
df = pd.DataFrame(portfolio_items)

# Check key columns (adjust based on your needs)
print(df[["name", "symbol", "share", "marketValue"]].head())

Fix 2: Use a Browser Renderer (If You Prefer HTML Parsing)

If you need to parse the fully rendered page (for example, if you want other elements too), use Selenium to simulate a browser that executes JavaScript.

Code with Selenium

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from bs4 import BeautifulSoup
import pandas as pd

URL = "https://www.wikifolio.com/de/de/w/wffalkinve"

# Initialize Chrome driver (make sure ChromeDriver is installed and in your PATH)
driver = webdriver.Chrome()
driver.get(URL)

try:
    # Wait up to 10 seconds for the portfolio rows to load
    WebDriverWait(driver, 10).until(
        EC.presence_of_element_located((By.CLASS_NAME, "c-portfolio"))
    )

    # Get the fully rendered page source
    soup = BeautifulSoup(driver.page_source, "lxml")

    # Now find all the portfolio rows
    portfolio_rows = soup.find_all("tr", class_="c-portfolio")

    # Extract data from each row
    portfolio_data = []
    for row in portfolio_rows:
        cells = row.find_all("td")
        row_content = [cell.get_text(strip=True) for cell in cells]
        portfolio_data.append(row_content)

    # Convert to DataFrame
    df = pd.DataFrame(portfolio_data)
    print(df.head())

finally:
    # Always close the browser
    driver.quit()

Notes for Selenium:

  • You need to install the correct driver for your browser (e.g., ChromeDriver for Chrome, GeckoDriver for Firefox) and ensure it matches your browser version.

Quick Tips to Avoid Future Headaches

  • Check for dynamic content first: If elements don’t show up in the initial HTML, they’re probably loaded via JS—use DevTools to find the underlying API.
  • Respect robots.txt: Always check https://www.wikifolio.com/robots.txt to confirm scraping the portfolio data is allowed.
  • Use proper headers: Adding a User-Agent header helps avoid being blocked by the server.

内容的提问来源于stack exchange,提问作者Julian

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:58:50