基于Python实现Windows新闻和兴趣组件或Microsoft Edge主页的交互需求及替代Pywinauto的方案咨询
Hey there! I get why Pywinauto feels clunky for this task—it's great for system-level UI automation, but overkill when you're dealing with web content inside Edge. Let's look at better, more focused alternatives that'll make scraping Edge's news recommendations and handling those interaction buttons a breeze.
1. Playwright (Top Pick)
Playwright is modern, browser-agnostic, and built specifically for web automation. It plays exceptionally well with Edge, handles element waiting automatically, and has a clean API that cuts down on boilerplate code. Here's how to use it for your use case:
Setup
First, install Playwright and the Edge browser binary:
pip install playwright playwright install msedge
Example Code
This script handles scraping article titles/links, extracting full article text, and interacting with buttons like "Hide this story":
from playwright.sync_api import sync_playwright with sync_playwright() as p: # Launch Edge (headless=False lets you see the browser in action) browser = p.chromium.launch(channel="msedge", headless=False) page = browser.new_page() # Navigate to Edge's home page (loads personalized news by default) page.goto("edge://home/") # Wait for news feed items to load (adjust selector if needed) page.wait_for_selector(".cfs-feed-item") # Scrape all article titles and links articles = page.query_selector_all(".cfs-feed-item") article_data = [] for article in articles: title = article.query_selector(".cfs-title").text_content().strip() link = article.query_selector("a").get_attribute("href") article_data.append({"title": title, "link": link}) print(f"Title: {title}\nLink: {link}\n") # Extract full text from the first article if article_data: page.goto(article_data[0]["link"]) # Wait for article body to load (adjust selector based on the site) page.wait_for_selector(".article-body") article_text = page.query_selector(".article-body").text_content().strip() print(f"--- Full Text of First Article ---\n{article_text}\n") # Example: Click "Hide this story" for the first article page.goto("edge://home/") # Open the article's options menu first_article_menu = page.query_selector(".cfs-feed-item .cfs-more-options") first_article_menu.click() # Click the hide button hide_button = page.query_selector(".cfs-option-hide") hide_button.click() browser.close()
Why Playwright Works Better
- It automatically waits for elements to be interactable, so you don't need to add manual sleep timers.
- It manages Edge browser binaries behind the scenes, no need to download matching drivers manually.
- Its API is intuitive and designed for web-specific tasks, unlike Pywinauto which targets system UI.
2. Selenium with EdgeDriver
If you're more familiar with Selenium's ecosystem, it's a solid alternative. Just note you'll need to manage EdgeDriver versions manually to match your Edge browser version.
Setup
pip install selenium
Download the matching EdgeDriver from Microsoft's official site (make sure it's in your system PATH).
Example Code
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC # Initialize Edge driver driver = webdriver.Edge() driver.get("edge://home/") # Wait for news feed to load wait = WebDriverWait(driver, 10) articles = wait.until(EC.presence_of_all_elements_located((By.CLASS_NAME, "cfs-feed-item"))) # Scrape article data article_data = [] for article in articles: title = article.find_element(By.CLASS_NAME, "cfs-title").text.strip() link = article.find_element(By.TAG_NAME, "a").get_attribute("href") article_data.append({"title": title, "link": link}) print(f"Title: {title}\nLink: {link}\n") # Extract full article text if article_data: driver.get(article_data[0]["link"]) article_text = wait.until(EC.presence_of_element_located((By.CLASS_NAME, "article-body"))).text.strip() print(f"--- Full Text of First Article ---\n{article_text}\n") # Click "Hide this story" driver.get("edge://home/") first_menu = wait.until(EC.element_to_be_clickable((By.CSS_SELECTOR, ".cfs-feed-item .cfs-more-options"))) first_menu.click() hide_button = wait.until(EC.element_to_be_clickable((By.CSS_SELECTOR, ".cfs-option-hide"))) hide_button.click() driver.quit()
3. Hybrid Approach (For News & Interest Sidebar)
If you specifically need to interact with the taskbar-based News & Interest sidebar (not just Edge's home page), you can combine a tiny bit of Pywinauto to open the sidebar, then use Playwright to handle the web content inside it:
- Launch Edge in debug mode first:
msedge.exe --remote-debugging-port=9222
- Use Pywinauto to click the News & Interest icon on the taskbar to open the sidebar.
- Connect Playwright to the running Edge instance to scrape and interact with the sidebar's web content:
from playwright.sync_api import sync_playwright with sync_playwright() as p: # Connect to the debug Edge instance browser = p.chromium.connect_over_cdp("http://localhost:9222") # Find the sidebar page (filter by URL containing "newsfeed") sidebar_page = next(page for page in browser.pages if "newsfeed" in page.url) # Now use sidebar_page to scrape articles, click buttons, etc.—same as the Playwright example above browser.close()
This way, you only use Pywinauto for the single system UI action, and handle all web-related tasks with Playwright's cleaner API.
内容的提问来源于stack exchange,提问作者Pradyumn Pathak

