How every number on every screen is computed
Written for a sceptical reader. If a figure appears anywhere in this product whose derivation is not on this page, treat that as a defect and not as a detail.
DRISHTI reads publicly available material about the Telangana government and reports what it finds. It is an evidence-grounded intelligence system, and it is deliberately not built to make anyone look good. Positive and negative findings are surfaced by the same rules, ranked by the same score, and given the same visual weight.
Nothing in it profiles individuals, infers political affiliation, scores persuadability, segments audiences or recommends messaging. Those capabilities are absent by design, not disabled by configuration.
The demo corpus is fixed, the clock is fixed at 2026-09-02 14:34:42.108+00, and every identifier is a UUIDv5 derived from content rather than generated randomly. Running the pipeline twice produces byte-identical output. No figure in this product changes on refresh, and none is produced by sampling.
Script is decided by counting glyphs in the Telugu, Latin, Devanagari and Arabic Unicode blocks. Where Telugu and Latin each exceed 15% of alphabetic characters, the record is classified mixed rather than forced into one language.
Latin text is ambiguous between English and Romanized Telugu, so it is tested against a list of 51 Telugu function words that do not occur in English. Two or more hits, or a marker density above 0.08, classifies the record as Telugu in Latin script. Confidence for that path is capped at 0.86 because it is a lexicon signal and not a model.
Entity linking runs in two passes. The first is an exact match against 194 curated aliases across 41 entities, longest alias first.
The second transliterates Telugu to Latin by a deterministic rule and reduces each word to a consonant skeleton — aspirates collapsed, nasals unified, vowels dropped after the first character, geminates collapsed last. “రేవంత్ రెడ్డి”, “Revanth Reddy” and “revanth reddi” all reduce to rvnt rd.
A skeleton match on its own never creates a link. It proposes a candidate, which is then confirmed against the curated alias table with a matching word count — otherwise “Kanal” would become “Kamal”.
21 topics, 376 weighted bilingual terms, lexicon version 1.0.0. Scores saturate as 1 − e^(−score / 2.2) so that ten repetitions of one term are not counted as ten pieces of evidence, then normalise across the topics that fired. Anything below 0.18 is discarded. Every classification records which terms matched, and those terms are shown on the record’s own page.
This is the part most media-monitoring products get wrong. Consider: “CM announces ₹500 crore relief for flood-hit farmers.” A single-axis sentiment model scores that positive or negative depending on which words it weights, and both answers are wrong, because the sentence contains two different facts.
- Event valence. Is the reported event good or bad news? Relief being sanctioned is positive.
- Author stance. Is the author supportive of or critical toward the subject? A reporter describing the announcement expressed no view at all.
- Audience signal. How the visible audience reacted, inferred from engagement shape where the platform exposes it. Frequently absent, and absent is reported as absent rather than as neutral.
- Subject sentiment. Sentiment directed at the tracked subject specifically, scoped to the sentence containing the mention. Null when the subject is not named.
A reporting outlet carrying a critical quotation is recorded as neutral reporting of criticism, not as criticism. Sarcasm signals never flip a label; they only reduce confidence.
Each record is reduced to a min-hash sketch of 5-word shingles. Jaccard similarity at or above 0.72 within 36 hours marks a near-duplicate; at or above 0.42 against a wire or official origin marks syndication. Exact content-hash matches are marked exact.
Independent origins, not mentions, drive every momentum figure in this product. Twenty outlets carrying one wire copy is one signal. Products that miss this overstate how much the public is saying, and systematically understate genuinely grassroots stories — which are exactly the ones a government most needs to see early. In the current corpus the duplicate rate is 0.0%.
Single-pass agglomeration over time-ordered records. Similarity is 0.46 × TF-IDF cosine + 0.26 × entity overlap + 0.16 × topic overlap + 0.12 × place overlap, linking at 0.42. The blend exists because a Telugu report and an English one about the same event share almost no tokens but do share entities, topics and places.
A cluster centroid stops absorbing after 12 records, a story cannot span more than 84 hours, and it closes after 30 hours of silence. Without those bounds the centroid drifts into a generic term soup that everything matches — in testing, one cluster absorbed 1,250 unrelated records before the caps were added.
Velocity is records per hour over the trailing 6 hours. Acceleration is (recent rate + 0.15) / (baseline rate + 0.15) against the preceding 12 hours. The Laplace smoothing prevents a jump from zero producing an infinite ratio.
Source diversity and platform spread are normalised Shannon entropy over the respective mixes. Reach is a proxy summed from source audience estimates and is never presented as a measurement of who actually saw something.
Escalation is multi-factor and every factor is named and signed, including the ones arguing against escalating. A system that lists only reasons to be alarmed is an alarm, and alarms get ignored.
Materiality is then capped by absolute volume: below 6 records the ceiling is emerging, below 15 developing, below 30 significant. Rate of change is the most useful signal available and the most easily fooled — two posts after a quiet period is an infinite percentage increase. Before this cap existed, every quiet topic that stirred slightly was escalating “34×” and the board filled with criticals, which is the same as having none.
A claim is an assertion with a magnitude that can be checked. Sentences are split into clauses first, the quantity is bound to the nearest predicate within 90 characters, and a unit is required.
Each of those three rules exists because its absence produced a confidently wrong verdict during development. Without clause splitting, “progress is 61 per cent, the date has moved to March 2027, and ₹28 crore is unaccounted for” bound “unaccounted” to “61 per cent” and produced a claim nobody had made. Without predicate binding, the parser reached for the first number in the sentence, which is very often a date. Without the unit requirement, a bare number became a checkable fact.
Claim identity is predicate | magnitude | unit | subject. Two statements of the same assertion, in either language and with any wording, collapse to one claim — which is what makes the restatement count meaningful rather than a word-frequency artefact.
Evidence is tiered by what a record is, never by whom it favours. A government press release is an official statement — a party’s account of itself — and never a primary document. An opposition statement is treated identically. Only signed orders, tabled figures and administrative records reach the primary tier.
Contradiction is asymmetric. A challenger must reach the standing of the claim’s own origin before it can carry the verdict: a signed departmental order is not refuted by an anonymous post asserting a different number. Two figures can only contradict each other if they share a unit — before that check existed, the verifier marked a primary record false because a different quantity in a different unit appeared nearby.
Restatement count contributes nothing to any verdict. A claim repeated a thousand times by accounts that all read one original post has been corroborated zero times. Circulation is reported because it says how far something travelled, and then explicitly excluded from the assessment.
Nothing is marked false for being unflattering. A figure failing verification says nothing about whether the concern attached to it is legitimate, and the two are ruled on separately.
District attribution comes from what a record says: 132 surface forms covering all 33 districts, major towns, GHMC localities and landmarks, in Telugu and English. A publisher’s home district is used only as a 0.4-confidence fallback and is excluded from every figure on Telangana Listens.
The location of a private individual is never inferred. Nothing is derived from an account, a device or a network.
Attention is reported per million residents alongside raw volume, using 2011 census figures re-apportioned to the 2016 district boundaries. Raw volume is a population map: without normalisation, Hyderabad is the brightest district on every screen and nothing anomalous is ever visible.
Findings and briefs are assembled clause by clause from computed values. There is no language model anywhere in the generation path. If a sentence cannot be built from a number the system holds, it is not written — which is why some sections say nothing rather than filling the space.
Every finding carries its blind spots, computed from the actual composition of the records behind it: geographic concentration of publishers, source concentration, dependence on official figures, share of low-confidence interpretations, and window length.
- Whether a conversation reflects public opinion. It reflects what was published on connected platforms by people who chose to post.
- How many people saw anything. Reach figures are proxies summed from source audience estimates.
- Causation. Findings say 'followed by' and 'coincided with'; nothing here establishes that an event caused a change in discussion.
- Anything about districts with no attributable records. Under-collection and quiet look identical from downstream.
- Whether a grievance is justified. VIGIL rules on assertions with magnitudes, not on whether a complaint is fair.
- What is being said in private, in closed groups, or on platforms that are not connected.
Every stage emits a provenance event for every object it touches, including the steps that found nothing — “no curated entity matched this text” is as much a part of the record as a match would be. The current corpus carries 7,087 such events.
Each records the agent, the stage, the method kind, a method reference, a version and a human-readable note. The evidence drawer on every screen is a direct rendering of them. A stage that skipped its provenance would make its own output unusable.