如何使用Python抓取DR新闻网站SENESTE NYT板块的推送数据?
Hey there! I’ve dug into DR Nyheder’s structure and figured out how to get that elusive "SENESTE NYT" content you’re after. Here’s a step-by-step breakdown:
1. Locate the Correct API Endpoint
The generic https://www.dr.dk/tjenester/newsapp-content/teasers endpoint you found pulls site-wide latest content, but the left-side "SENESTE NYT" section uses a collection-specific API built just for that feed. The right endpoint is:
https://www.dr.dk/tjenester/newsapp-content/teasers/collection/senestenyt?reqoffset=0&reqlimit=20
reqoffset: Controls pagination (start at 0 for the first batch of articles)reqlimit: Sets how many entries to fetch per request (max is typically around 100)
2. Verify the Response
When you hit this endpoint (make sure to mimic browser headers to avoid being blocked), you’ll get a JSON response with a teasers array. Each entry includes all the data you need for the section:
- Article headline (
headline) - Short summary (
lead) - Publication timestamp (
published) - Direct article URL (
url) - Thumbnail image details (
imageobject)
3. Mimic Browser Headers for Reliability
DR’s APIs might block requests that don’t look like they’re coming from a browser. Here’s an example using curl to test the endpoint properly:
curl "https://www.dr.dk/tjenester/newsapp-content/teasers/collection/senestenyt?reqoffset=0&reqlimit=20" \ -H "User-Agent: Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" \ -H "Referer: https://www.dr.dk/nyheder/"
4. Fallback: Check Inline Page Data
If the API ever stops working (e.g., temporary structural changes), check the page’s inline JavaScript. Search the page source for <script> blocks containing senestenyt — sometimes initial feed data is embedded directly in the page as JSON to speed up loading.
That should get you all the content from the left-side "SENESTE NYT" section! Let me know if you run into any snags.
内容的提问来源于stack exchange,提问作者bib

