如何从LinkedIn板块抓取帖子日期并获取指定文本的正确XPath?
Hey Jacek, let's work through how to correctly target that "2 dni temu" timestamp text from LinkedIn posts. The issue with your initial XPath is that the occludable-update ember-view element is just the top-level container for the post—your target timestamp is nested several layers deeper inside it, so we need to drill down to the right child element.
Here are a few reliable approaches using class-based targeting:
方案1:基于时间元素的特征类定位
LinkedIn consistently uses a set of small-text style classes for timestamps (like t-12, t-normal, t-black--light). We can combine these with the post container to grab the first post's timestamp:
//div[contains(@class, 'occludable-update ember-view')][1]//span[contains(@class, 't-12 t-normal t-black--light')]/text()
为什么这个有效:
contains(@class, ...)avoids issues with dynamic suffixes LinkedIn sometimes adds to class names (like random ember-view identifiers).- We first target the first post container with
[1], then traverse down to the span that holds the timestamp text.
方案2:精准匹配时间文本内容
Since you're targeting Polish-language timestamps ("dni temu" = "days ago"), we can add a text check to ensure we don't grab other secondary text in the post header:
//*[@class="occludable-update ember-view"][1]//*[contains(@class, 'update-components-actor__sub-description')]//span[contains(text(), 'dni temu')]/text()
额外提示:
- Always verify the current DOM structure using your browser's dev tools (F12): Right-click the timestamp, select "Inspect", then check the element's class names—LinkedIn occasionally updates these, so you might need to tweak the class strings.
- Avoid using absolute XPaths (like
/html/body/div[4]/...)—these break instantly if LinkedIn adjusts page layout. Stick to relative paths with class-based targeting. - If you're using a scraping tool (Selenium, Scrapy, etc.), make sure you're complying with LinkedIn's Terms of Service to avoid being blocked.
内容的提问来源于stack exchange,提问作者Jacek Antek

