957 records became 3 things to know
The value of this system is the reduction, and the reduction is only worth anything if you can audit it. Each step below names the rule that removed what it removed.
- 957records captured
Every item retrieved from a connected source, stored exactly as it arrived. Nothing at this layer has been modified, and nothing ever will be — corrections create new versions rather than overwriting.
- 197relevant to this deployment
Records concerning the Chief Minister, the Telangana government, its policies, services, or the issues put to it. A record can be highly relevant without naming the subject at all — a colony complaining about drains is squarely in scope.
760 records fell below the 0.35 relevance threshold and were retained but excluded from analysis.
- 197independent origins
Syndicated copies, near-duplicates and restatements removed. This is the number that describes how many people independently said something, and it is the number every momentum figure in this product is built on.
0 records (0%) were copies of something already counted.
- 92story clusters
Records grouped into events by TF-IDF similarity plus shared entities, topics and places — so a Telugu report and an English one about the same thing land together despite sharing almost no words.
- 16narrativesOpen →
Durable lines of public argument, assembled from clusters sharing a dominant topic and an anchor scheme, project or institution.
- 11material changesOpen →
Findings that reached developing or above. Everything below that level is recorded and visible, but does not claim a decision-maker's attention.
9 findings were retained at observation or emerging level.
- 3things to knowOpen →
What the morning brief leads with. Each one opens onto its arithmetic, its evidence, its counter-evidence and the original records.
Each one is retained, visible on its story cluster, and marked with the record it was found to duplicate and the similarity that produced the judgement. If the deduplication is wrong, it is wrong somewhere you can point at.
Kept in the store, excluded from narratives. A threshold this blunt will exclude things it should not; the alternative — no threshold — floods every measure with noise and is worse.
About 354 per finding. Every transformation on this page emitted one, including the steps that found nothing.
| Stage | Count | Share of raw | Rule applied |
|---|---|---|---|
| records captured | 957 | 100.0% | No exclusion at this step — records were grouped, not removed. |
| relevant to this deployment | 197 | 20.6% | 760 records fell below the 0.35 relevance threshold and were retained but excluded from analysis. |
| independent origins | 197 | 20.6% | 0 records (0%) were copies of something already counted. |
| story clusters | 92 | 9.6% | No exclusion at this step — records were grouped, not removed. |
| narratives | 16 | 1.7% | No exclusion at this step — records were grouped, not removed. |
| material changes | 11 | 1.1% | 9 findings were retained at observation or emerging level. |
| things to know | 3 | 0.31% | No exclusion at this step — records were grouped, not removed. |
Any of these figures can be walked back to the text that produced it. Pick a finding, open its evidence, open a story, open a record, and you arrive at the original Telugu, its translation, the URL, the timestamp and every step the pipeline took. The pipeline is inspectable too →