Playwright Python:如何仅筛选出指定格式(.jpg)的图片URL?
问题与解决方案
代码示例
import asyncio from playwright.async_api import Playwright, async_playwright, expect # 获取图片URL # 输出: img_urls.csv(包含{"property_id": str, "img_urls": []}格式的数据) async def run(playwright): image_urls = [] # 存储{"property_id": "value", "img_url": [img_urls]}格式的数据 browser = await playwright.chromium.launch(headless=False) context = await browser.new_context() # 打开新页面 page = await context.new_page() # 访问目标房产页面 await page.goto("https://www.zoopla.co.uk/for-sale/details/49240624/") # 点击"接受所有Cookie"按钮 await page.frame_locator("[aria-label=\"Privacy Manager window\\.\"]").locator("button:has-text(\"Accept all cookies\")").click() # 点击图片右箭头切换图片 for i in range(5): await page.locator("[data-testid=\"arrow_right\"]").click() # 获取所有图片URL imgs = await page.query_selector_all("img") for img in imgs: src = await img.get_attribute("src") print(src) # 关闭浏览器上下文和浏览器 await context.close() await browser.close() async def main() -> None: async with async_playwright() as playwright: await run(playwright) asyncio.run(main())
执行输出
上述代码执行后控制台输出如下:
https://lid.zoocdn.com/u/2400/1800/26d9845a91c7fe21834b531a292533dcf16f6754.jpg https://lid.zoocdn.com/u/2400/1800/2267900ffd5e795f568bf1305a5eab0b95e59e5f.jpg https://lid.zoocdn.com/u/2400/1800/75d0a22274ed94c1b33db52f5c5cc1022905df02.jpg https://lid.zoocdn.com/u/2400/1800/d459c1d0ff8e7a1b52f667a48e593f72f91ea368.jpg https://lid.zoocdn.com/u/2400/1800/0e21775e213c064fd83916b27e795536c816edb0.jpg https://maps.googleapis.com/maps/api/staticmap?size=792x398&format=jpg&scale=2¢er=51.535651,-0.006482&maptype=roadmap&zoom=15&channel=Lex-LDP&client=gme-zooplapropertygroup&sensor=false&markers=scale:2%7Cicon:https://r.zoocdn.com/assets/map-static-pin-purple-76.png%7C51.535651,-0.006482&signature=EZukT7ugiBKGYFT9F9phLleIXBs= https://lid.zoocdn.com/u/2400/1800/55e734ae277ee03b2d46f55298acd762727bd727.gif https://r.zoocdn.com/_next/static/images/natwest-dd532b27dc13112df4f05058c26a990a.svg https://st.zoocdn.com/zoopla_static_agent_logo_(584439).png
技术问题
每次点击图片右箭头切换图片后,控制台返回的URL数量会变化。想问下Playwright Python有没有官方的筛选方式,能直接选出以.jpg结尾或包含.jpg的图片URL,不用自己写Python代码判断?
解决方案
Playwright支持通过CSS选择器的属性匹配直接筛选符合条件的图片元素,这是官方原生支持的能力,无需后续在Python层面过滤,性能更优。具体有两种实现方式:
1. 筛选src属性以.jpg结尾的图片
使用CSS的$=属性匹配选择器,精准定位URL后缀为.jpg的图片:
# 替换原有的获取图片代码 jpg_imgs = await page.query_selector_all("img[src$='.jpg']") for img in jpg_imgs: src = await img.get_attribute("src") print(src)
2. 筛选src属性包含.jpg的图片
如果需要匹配URL任意位置包含.jpg的情况(比如示例中的Google静态地图URL),可以用CSS的*=属性匹配选择器:
# 替换原有的获取图片代码 contains_jpg_imgs = await page.query_selector_all("img[src*='.jpg']") for img in contains_jpg_imgs: src = await img.get_attribute("src") print(src)
优化后的完整代码
import asyncio from playwright.async_api import Playwright, async_playwright, expect async def run(playwright): browser = await playwright.chromium.launch(headless=False) context = await browser.new_context() page = await context.new_page() await page.goto("https://www.zoopla.co.uk/for-sale/details/49240624/") await page.frame_locator("[aria-label=\"Privacy Manager window\\.\"]").locator("button:has-text(\"Accept all cookies\")").click() for i in range(5): await page.locator("[data-testid=\"arrow_right\"]").click() # 筛选src以.jpg结尾的图片 print("=== 以.jpg结尾的图片URL ===") jpg_imgs = await page.query_selector_all("img[src$='.jpg']") for img in jpg_imgs: print(await img.get_attribute("src")) # 筛选src包含.jpg的图片 print("\n=== 包含.jpg的图片URL ===") contains_jpg_imgs = await page.query_selector_all("img[src*='.jpg']") for img in contains_jpg_imgs: print(await img.get_attribute("src")) await context.close() await browser.close() async def main() -> None: async with async_playwright() as playwright: await run(playwright) asyncio.run(main())
说明
- 每次切换图片后,重新执行上述选择器查询,就能获取当前页面符合条件的最新图片列表,适配元素数量变化的场景。
内容的提问来源于stack exchange,提问作者Andrea
相关产品推荐
相关产品推荐

