
Thousands of stories an hour is not a product. It is a problem. What turns the flood into something a person or an application can use is the information attached to each story about what it covers.
Why keywords are not enough
A keyword filter for “Apple” returns fruit prices. A filter for “merger” misses the story that says two firms “agreed to combine.” Keywords match letters. Readers care about meaning. Every product that relies on keywords alone ends up with a growing list of exceptions that someone has to maintain.
What a taxonomy adds
A taxonomy is a controlled list of things a story can be about, arranged so that broad terms contain narrower ones. Tagging a story with “renewable energy” also places it under “energy.” A reader who asked for the broad topic gets the story, and so does a reader who asked for the narrow one.
- Subjects: earnings, litigation, product launches, appointments.
- Industries, from sectors down to specialties.
- Geography: where it happened and whom it affects.
- Organizations, matched to a stable identifier and not just a name.
Consistency across sources
Publishers tag their own stories, but each uses a different scheme and many use none. The value of an aggregator’s taxonomy is that one set of tags applies to every source. A single filter then works across a national newspaper, a regional wire, a regulator’s announcements, and a specialist blog.
Tagging has to keep up
Tags applied an hour later are no help to a real-time feed. Automated classification does the work at speed. People maintain the scheme, review the difficult cases, and add terms as the world changes. New companies appear, industries split, and yesterday’s niche becomes a category.
How to judge a provider’s tagging
- Ask for the taxonomy itself and look at its depth in your field.
- Pick ten recent stories you know well and check the tags by hand.
- Look for missed stories as well as wrong ones. Missed stories are harder to see.
- Ask how often the scheme changes, and how you are told.


