--- name: estonian-news-media description: Estonian press. ERR in full; Postimees/Delfi/Õhtuleht headlines. Feeds hold 24-48h; older needs a general web search. module: uudised.mjs execution: get --- # Estonian News Media (ERR, Postimees, Delfi, Õhtuleht) The live layer of the Estonian press: what the outlets carry NOW, roughly the last 24–48 hours. For anything older, **Beyond the current window** below is the only route this repo has — the digitised-newspaper recipe that used to cover dated print was withdrawn on 2026-08-02 (see PLAN.md). The capability splits sharply, and honestly (all verified 2026-07-27): - **ERR** (public broadcaster, free): discovery AND full article text. - **Postimees, Delfi, Õhtuleht** (commercial, paywalled): discovery ONLY — headline, lead, time, link. ## Retrieve **Use `newsFeed()` and `errArticle()` (below).** One `newsFeed()` call is one fetch of one outlet, parsed, counted and newest-first; `newsOutlets()` lists what it knows. ```js const f = await newsFeed("err"); show(f.totals); // { items, newest, oldest } show(f.rows.slice(0, 3)); // title, link, published, category, lead const a = await errArticle(f.rows[0].link); ``` - `paywalled` on the result is the capability split, carried with the data: for Postimees, Delfi and Õhtuleht report headline, lead, time and link and hand over the link. `errArticle()` refuses those hosts by name rather than returning page furniture that reads like an article. - State counts from `totals.items`, and "the newest three" from the head of `rows` — every feed here is newest-first. - Topic questions inside the window: scan the returned titles, leads and keywords. No feed here has a search parameter. Anything older: **Beyond the current window** below. - ONE fetch per outlet per question. Postimees asks `Crawl-delay: 60` and disallows `/feed/`, `/rest/`, `/search`; Õhtuleht disallows every URL carrying a query string, so never append parameters there; Delfi disallows `/search` and `/otsing`. ERR imposes nothing on generic clients. ## Access Verified live **2026-08-01** — item counts are what the feeds returned that day. ERR (free, discovery AND full text): `https://www.err.ee/rss`, RSS 2.0, 50 items newest-first, each with `title`, `link`, `description` (the lead), `pubDate`, `category`. Rubric feeds have the same shape: `sport.err.ee/rss`, `kultuur.err.ee/rss`, `news.err.ee/rss` (English), `rus.err.ee/rss` (Russian). Article text is server-rendered on every `*.err.ee` host. The news sitemap (`https://www.err.ee/sitemap` → `sitemap/news0.xml`) carries the same stories with `news:keywords` when keywords help. Commercial, headline layer only: - Postimees `https://www.postimees.ee/rss` — 25 items; link is a `pmo.ee/` shortlink that redirects to the article. - Õhtuleht `https://www.ohtuleht.ee/rss` — 24 items; adds the author in `dc:creator`. - Delfi: NO RSS (`/rss` is 404). News sitemaps instead: `https://www.delfi.ee/sitemap/index.xml` → `https://www.delfi.ee/sitemap/news-1.xml` (newest, 50 URLs, Google News format with `news:keywords`) back to `news-20.xml`. Delfi stories also live on its rubric sites — `arileht.delfi.ee` (majandus), `epl.delfi.ee` (Eesti Päevaleht), `maaleht.delfi.ee`, `forte.delfi.ee`, `geenius.delfi.ee` — same publisher, same paywall, same headline-only boundary. If you ever read an ERR article by hand: take the `…` element, headline from its `

`, body from its `

` elements in order. Two traps `errArticle()` already handles — a ``, so the audio player's SVG arrives as the first "paragraph", and the body is HTML-escaped a second time, so every õäöü comes back as `ä` unless it is unescaped. ## Beyond the current window The feeds hold 1–2 days. Past that, this recipe's route IS a general web search, restricted to these domains — not a detour around the catalog but this guide's own instruction. The outlets' search surfaces are dead ends (probed 2026-07-27): ERR's `/otsing` and `/search` are one JS shell with no server-rendered results and its search API 404s; the commercial three robots-disallow their search paths (step 4). - `sites: ["err.ee"]` — or `["postimees.ee"]`, `["delfi.ee"]`, `["ohtuleht.ee"]`, several at once when the question is about "the press" — keeps the hits on the outlets instead of the open web. - The question names a period? Set `published_after` / `published_before` (YYYY-MM-DD) to it — "mida kirjutati juunis" is a date window, not a feed scroll. - The capability split above still rules the hits: an err.ee hit → fetch the article and read it in full (ERR article URLs stay live for years). A commercial hit → the search excerpt may be lead-only; report headline + date + link and hand over the link. Search improves DISCOVERY, not paywall access. - Digitised print (regional papers, anything pre-web) has NO source here since 2026-08-02. A web search reaches what is on the open web; a regional paper's pre-web archive is not, and the honest answer is to say so rather than to imply the press did not cover it. ## Return Per story: outlet, headline, publication time, link, plus rubric/author/keywords when the feed carries them. For ERR full text, quote or summarise only sentences present in your printed output, with the article's URL and pubDate. Say which outlets you actually checked — "the press" is a claim about four outlets, not one feed. ## Limits - **The commercial three end at the lead.** Their article pages are paywalled (PM+, Delfi tellijale, Õhtuleht premium): a plain fetch returns page furniture plus at most the lead, a teaser for premium stories. Never present a lead or teaser as the article, never state what a paywalled article says, never try to get around the wall — hand the citizen the link instead. - Feeds are a window, not an archive: ~50 items spans 17–48 h depending on outlet and news volume. Older coverage goes through **Beyond the current window** above. - Absence from a feed is not absence from the press — regional papers, broadcast and magazines are not here. - No reach or impact claims: a feed says what was published, not what was read or believed. ## Verify - A feed counts as retrieved only when the payload parses as XML and yields at least one item with all three of title, link and date (`` with ``, or sitemap `` with ``). An HTML page, a bot-challenge page or an empty shell is not a feed — `newsFeed()` throws on one rather than returning an empty result, and the answer is to report the retrieval as failed, not to work around it. - An **empty** `

` extraction is a retrieval failure, not an empty article, and `errArticle()` throws on it. **One** paragraph is neither: ERR publishes broadcast listings whose whole body is a presenter's name, and the result flags that in `note` instead of failing a working page. - Every headline, time and link stated must appear in printed output. Route self-check, re-verified live 2026-08-01 through the module: `newsFeed("err")` returned 50 items, `newsFeed("postimees")` 25, `newsFeed("ohtuleht")` 24, `newsFeed("delfi")` 50 sitemap URLs, and `errArticle("https://www.err.ee/1610087302/valitsus-ei-toeta-jahtumisaja-kehtestamist-korgetele-ametnikele")` yielded 16 body paragraphs (3,265 chars), published 2026-07-27, editor Märten Hallismaa. ## Module API newsFeed(outlet) -> {source, page, outlet, paywalled, totals, rows, note} errArticle(url) -> {source, title, published, editor, totals, paragraphs, text, note}