如何使用pywinauto从谷歌Chrome弹窗中提取"Assunto"列的列表文本?
Hey there! I see you're trying to extract the "Assunto" column text from a Chrome pop-up using pywinauto, and you've only been able to identify the top-level Chrome rendering widgets so far. Let's break down why this is happening and how to fix it.
The core issue here is that Chrome renders web content in a proprietary layer (the Chrome_RenderWidgetHostHWND controls you're seeing). These are single container controls that don't expose individual web elements like table columns as standard Windows UI components—so pywinauto can't directly "see" the "Assunto" column through its native control scanning.
Here are three actionable solutions tailored to your setup (Pywinauto 0.6.8, Python 3.10.7, Windows 10):
1. Use Selenium (Recommended Approach)
Selenium is built specifically for interacting with web browsers and can directly access DOM elements like table columns. This is the most reliable method for web content extraction.
First, install Selenium and set up ChromeDriver (matching your Chrome version):
pip install selenium webdriver-manager
Then use this code to extract the "Assunto" column:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.chrome.service import Service from webdriver_manager.chrome import ChromeDriverManager import time # Initialize Chrome driver (automatically installs compatible ChromeDriver) driver = webdriver.Chrome(service=Service(ChromeDriverManager().install())) # Navigate to your AgênciaNet page (replace with the actual URL) driver.get("https://your-agencia-net-url-here") # Wait for the page/table to load driver.implicitly_wait(10) time.sleep(1) # Extract the "Assunto" column (adjust the XPath selector to match your table's structure) # Example: Find all cells under the "Assunto" header column assunto_cells = driver.find_elements(By.XPATH, "//table//th[text()='Assunto']/parent::tr/following-sibling::tr/td[position()=2]") # Print each item in the column for cell in assunto_cells: print(cell.text) # Clean up driver.quit()
2. Pywinauto + Clipboard (Quick, No Extra Tools)
If you can't use Selenium, you can simulate selecting and copying the table content, then parse the clipboard data. This works best for simple, well-formatted tables.
import pyperclip from pywinauto import Application import time app = Application().connect(title_re='AgênciaNet', timeout=10) time.sleep(1) window = app.Chrome_WidgetWin_1 window.set_focus() window.maximize() # Simulate selecting all content and copying to clipboard window.type_keys('^a') # Ctrl+A to select window.type_keys('^c') # Ctrl+C to copy time.sleep(0.5) # Wait for clipboard to populate # Get clipboard content and parse it clipboard_text = pyperclip.paste() lines = clipboard_text.split('\n') # Extract the "Assunto" column (adjust column index based on your table's formatting) for line in lines: columns = line.split('\t') # Assumes tab-separated table; use ',' for CSV if needed if len(columns) >= 2: print(columns[1]) # Replace index 1 with the actual column position of "Assunto"
3. Pywinauto + UIAutomation Library
Chrome supports UI Automation, so you can use the uiautomation library to dig into the rendered web content. This is more manual but works if you need to stick with Windows UI control methods.
First install the library:
pip install uiautomation
Then use this code to locate the column:
import uiautomation as auto from pywinauto import Application import time app = Application().connect(title_re='AgênciaNet', timeout=10) time.sleep(1) window = app.Chrome_WidgetWin_1 window.set_focus() # Locate the Chrome window via UIAutomation chrome_window = auto.WindowControl( searchDepth=1, Name='AgênciaNet - Secretaria de Economia do Distrito Federal - Google Chrome' ) # Navigate to the web content container (you may need to tweak this based on your page's UI tree) web_pane = chrome_window.FindControl(auto.ControlType.ControlTypePane, Name='Chrome Legacy Window') table = web_pane.FindControl(auto.ControlType.ControlTypeTable) # Find the "Assunto" header and its corresponding column cells assunto_header = table.FindControl(auto.ControlType.ControlTypeHeaderItem, Name='Assunto') assunto_column_rows = assunto_header.GetParent().FindChildren(auto.ControlType.ControlTypeDataItem) # Print each cell's text for row in assunto_column_rows: print(row.Name)
内容的提问来源于stack exchange,提问作者Koba

