All News. All Sources.

How PageNews Works

Last updated

This page describes exactly how a headline gets from a publisher's feed onto PageNews, and how the ordering on each page is decided. Where a process is automated and imperfect, we say so.

1. Collection

A scheduled job runs every fifteen minutes and requests the RSS or Atom feed of each active source. Feeds are fetched with conditional requests, so a publisher that has posted nothing new returns a small “not modified” response rather than the whole document.

2. Normalisation

For each item we keep the headline verbatim, the publisher's own summary, the publication timestamp, the image the feed supplies if any, and the link. Timestamps are converted to UTC. Campaign and tracking parameters are stripped from links. Where a publisher dates an item in the future, usually a timezone mistake on their side, we clamp it to the current time so it cannot sit permanently at the top of a list.

3. Pictures

When a feed supplies an image we make a small thumbnail from it and serve that copy ourselves. Publishers’ own images range from three kilobytes to several megabytes; loading them directly would make a page of forty headlines enormous. The thumbnails are deleted when the story passes out of the retention window. Headlines without a picture simply do not show one — nothing is substituted.

4. Removing duplicates

The same article often appears under several URLs. We identify an article by a hash of its address once tracking parameters, www. and trailing slashes are removed, so those collapse into one entry.

5. Grouping stories

When several publishers report the same event, we group their coverage so the event appears once with all its sources attached, rather than as separate near-identical rows. Grouping is automatic and based on headline similarity and timing. It is not perfect: closely related but distinct events can occasionally be merged, and coverage that uses very different wording can be missed.

6. Categories

Each story is assigned one primary section. Assignment is automatic and based on the content of the headline and summary rather than on which feed it arrived through, because publishers' own feed categories are broad and frequently wrong for an individual item. Every classification carries a confidence value; when confidence is too low, the story stays in Latest rather than being filed confidently into the wrong section.

7. Trending

Trending topics are computed from the names and phrases that appear across headlines published in the last twenty-four hours, weighted towards terms whose frequency has risen recently rather than towards terms that are simply always common. A topic needs coverage from more than one publisher before it appears. Trending measures what publishers are writing about.

8. Most Read

Most Read counts how many times PageNews visitors clicked through to a story, over a recent window. It measures reader behaviour on this site only, and nothing else. Known crawlers are excluded and repeated clicks from the same visitor within a short period count once. On a quiet day the numbers involved are small, and we do not dress them up as anything larger.

Trending and Most Read are deliberately separate: one reflects publishers, the other reflects readers.

9. Ordering

Top Stories combines recency, how many independent publishers are covering the event, and reader interest. It is not a pure reverse-chronological list. We also avoid letting one publisher occupy several consecutive slots purely because it posts frequently — though a genuinely major story is never demoted for the sake of variety.

Nothing here is paid

No publisher pays to be included, to rank higher, or to appear in Top Stories. There is no advertising relationship affecting the ordering of any page on this site.

When it gets things wrong

Automatic classification and grouping make mistakes. If you find a story in the wrong section, two unrelated events merged together, or a headline attributed to the wrong publisher, please tell us through the corrections page.